In this section

Predicting Which Generated Answers Will Be Wrong

Module 0

Introduction

Checking a generated answer costs you thirty seconds every time and finds nothing on most of them. That is why it gets abandoned under load, and why a habit built on checking alone does not survive a busy fortnight. This section is about the cheaper move: reading your own request and predicting, before any output exists, whether this is one of the answers likely to be wrong.

Five conditions do most of that work, and all five are properties of what you asked rather than of what came back. Three of them can be removed by rewriting the request instead of checking the reply, which prevents the failure rather than catching it. By the end you will have a pre-flight you can run in about thirty seconds and a rule for how much scrutiny an artifact has earned.

Scenario

Eleven alerts, four generated artifacts in the last twenty minutes, and roughly a minute each. You cannot run a full check on all four and you should not run none. What you need is a way to tell, before reading a single line of output, which of the four is most likely to be wrong and which check will settle it.

01

Two calibrations, both yours

What the tool is likely to get wrong, and what you are

The previous sub established that you cannot ask a model how confident it is and get an answer worth having. That leaves two calibrations, and both belong to you.

Calibration about the task. Before you delegate, how likely is this particular request to produce a wrong answer? This is predictable from properties you can see, and the next section lists them.

Calibration about yourself. How often does your own judgment of a generated artifact turn out to be right? Most analysts have never measured this and most assume it is higher than it is.

The two combine into the only reliable practice available: knowing which requests are risky, and knowing how much your own review is worth on the risky ones.

TWO CALIBRATIONS, AND THEY ARE MEASURED DIFFERENTLY
 
ABOUT THE TASK      how likely is THIS request to go wrong?
  measured from     properties of your own sentence
  measured when     before you send it
  costs             about five seconds
 
ABOUT YOURSELF      how often is my judgment of an artifact right?
  measured from     scoring five artifacts before you check them
  measured when     over a working week
  costs             nothing you were not already doing
 
NEITHER OF THESE IS A PROPERTY OF THE MODEL.

That is the whole reason both are worth acquiring. Nothing about the tool changes either number, so neither goes stale when the tool does, and both are yours to improve.

02

Five conditions that predict trouble

Each one in the request, mapped to the failure it produces

Every one of these is visible before you read the answer.

1. Your request contained a relative time. "Around the alert", "recently", "in the last few days". The conversion to an absolute window is invisible in the output and an off-by-one-day window still returns rows. This one condition accounts for most silent windows.

2. The question was about your estate rather than about the world. "Is this suspicious", "is this normal". Unanswerable, and answered anyway in generically true terms.

3. The answer is a negative or a zero. No evidence, no other hosts, zero rows. Absence cannot be established from a sample.

Conditions 1, 2 and 5 are removed by rewriting the request. 3 and 4 cannot be, and those are the two that earn a check.

4. The artifact combines sources. Any join or "cross-reference these two". The combining step is where key uniqueness matters and where the error hides inside a count.

5. The output contains a specific figure you did not calculate. A count, a duration, a percentage, a ranking. Generated rather than measured.

SCORE THESE FIVE BEFORE YOU SCROLL. WHICH CONDITIONS FIRE?
 
  a  "pull the sign-ins for this account around the alert"
  b  "cross-reference these alerts against the device records"
  c  "is it normal for this service account to hit a laptop?"
  d  "how many files did the attacker touch?"
  e  "show me every sign-in for r.scott between 18:00Z and 06:00Z"
 
ONE OF THE FIVE FIRES NOTHING AT ALL.

Why these are worth more than a confidence score. They are properties of the request and the artifact, not claims by the system about itself. You can check them without trusting anything, they are visible in under ten seconds, and they map one-to-one onto the six failure modes. A confidence score is a sentence the model produced; these are facts about what you asked for.

READ THE REQUEST, PREDICT THE FAILURE IN WHAT YOU ASKED WHAT IT PRODUCES A relative time: "around the alert" 2 Silent window A question about your estate A generically true answer An answer that is a negative or zero 4 Confident absence An artifact combining two sources 3 Wrong join key A figure you did not calculate 6 Invented precision

Every condition is a property of what you asked, so you can read all five before any output exists.

03

Knowing your own hit rate

The number that tells you where your attention belongs

The second calibration is uncomfortable and it is worth acquiring, because every decision about how much to verify depends on it.

The method is simple. Over a week, when you review a generated artifact, write down your judgment before you check it: sound, or unsound and why. Then check properly. Count.

Most people find three things, and they are consistent enough to predict.

THE SCORECARD, ONE WEEK
 
  artifact   your call     actual     
  ────────────────────────────────────
  1          sound         sound      
  2          sound         UNSOUND    <- the one that costs you
  3          unsound       unsound    
  4          unsound       SOUND      <- the one that costs the tool
  5          sound         sound      
 
TWO NUMBERS COME OUT OF THIS, NOT ONE: WHAT YOU MISSED, AND
WHAT YOU FLAGGED THAT WAS FINE.

Your hit rate is lower than you expect on plausible artifacts. Everyone catches the obviously broken ones. The interesting number is your rate on the ones that look fine, and it is lower than it feels, because the cases you remember are the ones you caught.

Your false-positive rate is not zero. You will flag correct artifacts as suspicious, and an analyst who distrusts everything has stopped getting value from the tool.

Your rate varies by category, and predictably. You are good at spotting wrong queries in tables you know well and poor in tables you use rarely, which is the same estate-knowledge boundary from the previous sub applied to yourself.

WHAT A WEEK OF SCORING USUALLY SHOWS
 
  obviously broken artifacts     you catch nearly all
  plausible broken artifacts     you catch far fewer
  sound artifacts                you flag some of them
  tables you use daily           you are good
  tables you touch monthly       you are not
 
THE LAST TWO ROWS ARE WHERE THE THIRTY SECONDS SHOULD GO.

What to do with the number. Not to feel bad about it. To allocate. If you catch most errors in identity queries and few in endpoint ones, the endpoint queries are where the extra thirty seconds goes.

04

The fluency trap

Why this gets harder as the tooling improves

One effect deserves separating out, because it operates on your calibration rather than on the artifact.

A response with clear structure, hedging in the right places and a confident closing sentence reads as more reliable than a terse one, and none of those properties is evidence about the claim.

A better-written answer is judged more likely to be correct, and the two are unrelated.

The effect is worth naming because it is not a mistake you can decide to stop making. It operates on the reading, before any judgment you would recognize as a judgment, which is why noticing that an answer was well-argued is the moment to slow down rather than the moment to accept.

Two answers to one question · assistant output

A. No, that pattern is not typical. Service accounts of this kind are normally scoped to server workloads, and interactive-style authentication to an end-user device would be unusual in most environments, though it can occur where legacy applications require it. I would treat it as worth reviewing.

B. I cannot tell. That depends on how your service accounts are scoped, which is not something I have.

B is correct and A is the one that reads like expertise. A hedges in the right places, names a legitimate exception, and closes with a recommendation, and every one of those properties is a property of the writing rather than of the claim.

This matters specifically because fluency is increasing. Each generation of these systems produces better-organized, better-hedged, more professional-sounding output, and none of that improvement is in the direction of being right about your estate. The signal that feels like reliability is getting stronger while its correlation with reliability stays where it was.

THE FLUENCY TRAP each model generation how RELIABLE it reads how reliable it IS about your estate The signal that feels like reliability strengthens. What it tracks does not.

The rising line is the one you feel. The flat one is the one that decides whether the answer is right about your estate.

The practical form of this is worth recognizing in yourself. If your reason for accepting an answer is that it was well-argued, you have been persuaded rather than convinced, and the distinction is that a persuasive answer changes your belief without providing evidence.

05

The five conditions against a real request

Scoring a request that reads as completely reasonable, before any output exists

Take a request an analyst made on the Northgate brute force and score it before reading the answer.

THE REQUEST, AS IT WAS TYPED
 
  "The alert on r.scott fired around 23:00 on the 2nd. Show me what
   happened with that account recently, and tell me if it looks
   like the attack worked."
 
SCORE IT AGAINST THE FIVE BEFORE YOU READ ON.

Condition 1 fires twice. "Around 23:00" and "recently" are both relative, and the second has no anchor at all. Any window that misses 22:14 to 22:55 excludes the attack and still returns ordinary sign-ins.

Condition 2 fires. "Tell me if it looks like the attack worked" is a judgment about this estate. The answer will be about what successful brute forces generally look like.

Condition 3 is live. "The attack does not appear to have succeeded" is a negative claim, and it means the query it wrote surfaced no success rather than that none occurred.

Condition 4 is absent. No correlation across tables is required.

Condition 5 is likely. Any summary of "what happened with that account" will contain counts, and those counts will be generated.

SCORED, BEFORE ANY OUTPUT EXISTS
 
1  relative time        FIRES TWICE   "around 23:00", "recently"
2  about your estate    FIRES         "does it look like it worked"
3  answer could be      LIVE          a "no evidence" answer is a
   a negative                         claim about the query, not
                                      about the account
4  two sources          absent        no correlation needed
5  a figure to repeat   LIKELY        any summary carries counts
 
FOUR OF FIVE. NOTHING ABOUT THE REQUEST IS BADLY PHRASED.

Four of five, on a request that reads as completely reasonable. That is the point of the list: this is not a badly phrased request, it is how people actually ask, and the conditions fire on ordinary language rather than on carelessness.

The rewritten version costs nothing.

THE SAME QUESTION, ASKED SO THE CONDITIONS DO NOT FIRE
 
  "Show me every authentication for r.scott@ne.com between
   2026-03-02T18:00Z and 2026-03-03T06:00Z, including result
   codes, sorted by time."
 
ABSOLUTE WINDOW. NO JUDGMENT REQUESTED. NO FIGURE TO GENERATE.
CONDITIONS 1, 2 AND 5 ARE GONE RATHER THAN CHECKED.

It produces the evidence the original question wanted and took five seconds longer to type.

Scoring the request rather than the answer moves the check earlier, where it is cheaper. Analysts who do this well are not more suspicious than everyone else. They are suspicious at a different point.

Recalibrating when the tool changes

Every calibration in this sub is about a specific tool at a specific time. When the model behind it changes, and it will without announcement, your hit rate and your sense of which requests are risky are both measured against a system that no longer exists.

WHAT AN UPGRADE CHANGES, AND WHAT IT DOES NOT
 
  the six failure modes        unchanged, they follow from
                               the mechanism
 
  which of the six you meet    CHANGES, and without notice
 
  your hit rate                measured against a system that
                               no longer exists
 
  how well-written it reads    IMPROVES, which is the trap

The failure modes do not change, because they follow from the mechanism. The distribution across them does. A new model might be better at time windows and no better at negative claims, which moves where your attention belongs without changing anything you would notice casually.

The cheap response is to re-run the twenty-request evaluation from the previous sub after any significant change, and to treat a sudden improvement in fluency as a reason to check more rather than less. An answer that got noticeably better-written did not necessarily get more correct, and the fluency trap is strongest immediately after an upgrade.

Test it on your own assistant

Try this Give it a relative time and watch what it does
An alert fired around 11pm last night on one of our user accounts.

Write me a KQL query against SigninLogs to show what happened with that account around the alert.

What to look at. The time filter. You gave it "around 11pm last night" and "around the alert", and it has to turn that into something absolute. Look at what window it chose, and whether the query tells you which window it chose.

What this demonstrates. Something will come back with a specific time range in it, invented from nothing, and the query will run against real data and return real rows. That is the silent window at its source.

Now rewrite the request with an absolute range, between (datetime("2026-03-02T18:00:00Z") .. datetime("2026-03-03T06:00:00Z")), and notice that the failure is gone rather than caught. Five seconds of rewriting removed a condition that would otherwise have needed a check.

06

The pre-flight

Reading your own request back, and which conditions a rewrite removes

The five conditions are only worth anything if you can run them at speed, and they run against your own sentence rather than against anything that came back.

Pre-flight Five things to look for in your own sentence, before you send it
READ YOUR OWN REQUEST BACK
 
1. RELATIVE TIME?       "around the alert", "recently", "lately"
                        -> silent window. REWRITE IT, do not check it.
 
2. ABOUT YOUR ESTATE?   "is this normal here", "does this look
                        suspicious" -> you will get a general truth.
 
3. COULD THE ANSWER     a zero, a "no evidence of", an absence
   BE A NEGATIVE?       -> confident absence. Prove it can return rows.
 
4. TWO SOURCES?         anything correlating tables -> wrong join key.
 
5. WILL YOU REPEAT      any count, duration or total you will put in
   A FIGURE?            a ticket -> invented precision.
 
CONDITIONS 1, 2 AND 5 ARE REMOVED BY REWRITING THE REQUEST.
3 AND 4 ARE NOT, AND THOSE ARE THE ONES THAT EARN A CHECK.
Three of the five you can remove before any output exists. That is the half of the work that costs five seconds rather than thirty.

Copy that somewhere you will see it. The checks you run on the output are in 0.3, one per failure mode. This card is the half that runs before anything has been generated.

The obvious objection is that plenty of requests are not yours. The card runs backwards: reconstruct the request that would have produced the artifact, then score that. A summary asserting what happened "recently" came from a request containing a relative time, whether or not you can see the sentence.

The three phrases that cost the most

Two of the five conditions are triggered by how the request was worded rather than by anything the model did. Three phrases account for most of it, and all three are how people naturally speak.

Rewrite card The three phrasings that cause most of the damage
YOU SAY                  YOU GET                      SAY INSTEAD
─────────────────────────────────────────────────────────────────────
"around the alert"       a window that may miss it    "between 18:00Z
                                                       and 06:00Z"
 
"recently" / "lately"    an anchor from today, not    an absolute date
                         from the incident             range
 
"is this suspicious?"    a true statement about       "show me every X
                         accounts in general           for this account"
Rewriting the request is cheaper than checking the answer, and it removes the failure rather than catching it.

Rewriting the request is cheaper than checking the answer, and it removes the failure rather than catching it.

07

A working heuristic

Green, amber and red, and what each earns

Everything above compresses into something usable at speed.

Green. Translation, no relative time, no join, no figure to repeat. Read it against your intent and use it. This is most requests and it should feel fast.

Amber. Any of the five conditions present. Run the specific check that condition implies, not a general review.

Red is not a level of care. It is an instruction to go and get the evidence yourself rather than to check somebody else's.

Red. The request was about your estate, or the answer is going into a report, or the decision that follows is expensive to reverse. Do not verify the artifact. Go and get the evidence yourself, and use the generated answer as a hypothesis to test rather than a result to check.

HOW MUCH CHECKING DOES THIS ARTIFACT EARN? GREEN Translation, absolute time, no join, no figure to repeat read and use Most requests. This should feel fast. AMBER Any of the five conditions is present in the request one named check Thirty seconds, not a general review. RED Estate question, or going into a report, or hard to undo get the evidence Do not verify the artifact. Treat it as a hypothesis. Red is not "the important alerts". The alert that turns out to matter rarely announces itself in advance.

Most artifacts are green. A habit that treats everything as amber costs more than it returns and gets dropped.

The bands are a way of deciding how much attention an artifact earns, not how serious the alert is. Getting that distinction right is what keeps the habit affordable, because most of what you see is green and treating it as amber is how the practice dies.

The Reasonable Mistake

The category people get wrong

Red is not "the important alerts". It is defined by whether the answer depends on your estate or is expensive to reverse, and plenty of routine alerts qualify while plenty of dramatic ones do not. An analyst who reserves care for incidents that feel serious will apply it to the wrong ones, because the alert that turns out to matter rarely announces itself in advance.

So the band is read off the request and the consequence, never off how the alert feels. Both of those you can name in a sentence before you have looked at anything, which is what makes the classification cheap enough to actually perform.

08

Why predicting beats checking

Where the cost sits, and why one of the two habits survives a shift

Prediction and checking are not two names for the same habit, and the difference is where the cost sits.

Checking happens after an artifact exists. It costs thirty seconds each time, it competes with everything else on your queue, and it is the step that goes first under pressure, because the reward for doing it is almost always nothing.

Predicting happens before, from your own sentence, and it costs about five seconds. It also changes what you do next rather than only what you conclude: a request exposed to a silent window is a request you rewrite with an absolute window, and the failure never occurs.

You have not caught an error. You have prevented one, and that difference compounds across a shift in a way catching does not.

That is also why the five conditions are properties of the REQUEST rather than of the output. Anything you can only assess after reading the answer arrives too late to change the question, and by then you are in the expensive half of the work.

                      CHECKING            PREDICTING
 
runs                  after the artifact  before the request
                      exists              is sent
 
costs                 ~30 seconds         ~5 seconds
 
finds                 nothing, most       nothing, most times
                      times
 
outcome               you caught it       it never happened
 
what goes first       this one, because   nothing. It is part of
under pressure        the reward is       typing the request
                      almost always
                      nothing

The bottom row is the one that decides it. A habit that competes with your queue loses to your queue, and the only version of this that survives a working month is the version that is part of writing the request rather than a step you perform afterwards.

09

Practice

Rewrite three requests before sending them
hands on

You cannot calibrate against somebody else's hit rate. Five artifacts of your own, scored before you check them, is the smallest sample that tells you which half of the problem you actually have.

Practice Rewrite the request, do not check the answer
  1. Catch yourself about to type "around the alert" and give an absolute time range instead.
  2. Catch "recently" and give an absolute date range.
  3. Catch "is this normal?" and ask "how many times in ninety days?" instead.
Rewriting removes the failure. Checking only catches it after it has happened.

Count how many times in one shift you were about to use one of the three. The number is usually higher than people expect, and it is the cheapest habit in this module because it costs five seconds and happens before anything has gone wrong.

Next: section 0.6 deals with the one risk that has nothing to do with correctness, which is what leaves your estate when you ask for help.