Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
Predicting Which Generated Answers Will Be Wrong
Introduction
Checking a generated answer costs you thirty seconds every time and finds nothing on most of them. That is why it gets abandoned under load, and why a habit built on checking alone does not survive a busy fortnight. This section is about the cheaper move: reading your own request and predicting, before any output exists, whether this is one of the answers likely to be wrong.
Five conditions do most of that work, and all five are properties of what you asked rather than of what came back. Three of them can be removed by rewriting the request instead of checking the reply, which prevents the failure rather than catching it. By the end you will have a pre-flight you can run in about thirty seconds and a rule for how much scrutiny an artifact has earned.
Scenario
Eleven alerts, four generated artifacts in the last twenty minutes, and roughly a minute each. You cannot run a full check on all four and you should not run none. What you need is a way to tell, before reading a single line of output, which of the four is most likely to be wrong and which check will settle it.
Two calibrations, both yours
What the tool is likely to get wrong, and what you areThe previous sub established that you cannot ask a model how confident it is and get an answer worth having. That leaves two calibrations, and both belong to you.
Calibration about the task. Before you delegate, how likely is this particular request to produce a wrong answer? This is predictable from properties you can see, and the next section lists them.
Calibration about yourself. How often does your own judgment of a generated artifact turn out to be right? Most analysts have never measured this and most assume it is higher than it is.
The two combine into the only reliable practice available: knowing which requests are risky, and knowing how much your own review is worth on the risky ones.
TWO CALIBRATIONS, AND THEY ARE MEASURED DIFFERENTLY
ABOUT THE TASK how likely is THIS request to go wrong?
measured from properties of your own sentence
measured when before you send it
costs about five seconds
ABOUT YOURSELF how often is my judgment of an artifact right?
measured from scoring five artifacts before you check them
measured when over a working week
costs nothing you were not already doing
NEITHER OF THESE IS A PROPERTY OF THE MODEL.
That is the whole reason both are worth acquiring. Nothing about the tool changes either number, so neither goes stale when the tool does, and both are yours to improve.
Five conditions that predict trouble
Each one in the request, mapped to the failure it producesEvery one of these is visible before you read the answer.
1. Your request contained a relative time. "Around the alert", "recently", "in the last few days". The conversion to an absolute window is invisible in the output and an off-by-one-day window still returns rows. This one condition accounts for most silent windows.
2. The question was about your estate rather than about the world. "Is this suspicious", "is this normal". Unanswerable, and answered anyway in generically true terms.
3. The answer is a negative or a zero. No evidence, no other hosts, zero rows. Absence cannot be established from a sample.
4. The artifact combines sources. Any join or "cross-reference these two". The combining step is where key uniqueness matters and where the error hides inside a count.
5. The output contains a specific figure you did not calculate. A count, a duration, a percentage, a ranking. Generated rather than measured.
SCORE THESE FIVE BEFORE YOU SCROLL. WHICH CONDITIONS FIRE?
a "pull the sign-ins for this account around the alert"
b "cross-reference these alerts against the device records"
c "is it normal for this service account to hit a laptop?"
d "how many files did the attacker touch?"
e "show me every sign-in for r.scott between 18:00Z and 06:00Z"
ONE OF THE FIVE FIRES NOTHING AT ALL.
Why these are worth more than a confidence score. They are properties of the request and the artifact, not claims by the system about itself. You can check them without trusting anything, they are visible in under ten seconds, and they map one-to-one onto the six failure modes. A confidence score is a sentence the model produced; these are facts about what you asked for.
Every condition is a property of what you asked, so you can read all five before any output exists.
Knowing your own hit rate
The number that tells you where your attention belongsThe second calibration is uncomfortable and it is worth acquiring, because every decision about how much to verify depends on it.
The method is simple. Over a week, when you review a generated artifact, write down your judgment before you check it: sound, or unsound and why. Then check properly. Count.
Most people find three things, and they are consistent enough to predict.
THE SCORECARD, ONE WEEK
artifact your call actual
────────────────────────────────────
1 sound sound
2 sound UNSOUND <- the one that costs you
3 unsound unsound
4 unsound SOUND <- the one that costs the tool
5 sound sound
TWO NUMBERS COME OUT OF THIS, NOT ONE: WHAT YOU MISSED, AND
WHAT YOU FLAGGED THAT WAS FINE.
Your hit rate is lower than you expect on plausible artifacts. Everyone catches the obviously broken ones. The interesting number is your rate on the ones that look fine, and it is lower than it feels, because the cases you remember are the ones you caught.
Your false-positive rate is not zero. You will flag correct artifacts as suspicious, and an analyst who distrusts everything has stopped getting value from the tool.
Your rate varies by category, and predictably. You are good at spotting wrong queries in tables you know well and poor in tables you use rarely, which is the same estate-knowledge boundary from the previous sub applied to yourself.
WHAT A WEEK OF SCORING USUALLY SHOWS
obviously broken artifacts you catch nearly all
plausible broken artifacts you catch far fewer
sound artifacts you flag some of them
tables you use daily you are good
tables you touch monthly you are not
THE LAST TWO ROWS ARE WHERE THE THIRTY SECONDS SHOULD GO.
What to do with the number. Not to feel bad about it. To allocate. If you catch most errors in identity queries and few in endpoint ones, the endpoint queries are where the extra thirty seconds goes.
The fluency trap
Why this gets harder as the tooling improvesOne effect deserves separating out, because it operates on your calibration rather than on the artifact.
A response with clear structure, hedging in the right places and a confident closing sentence reads as more reliable than a terse one, and none of those properties is evidence about the claim.
The effect is worth naming because it is not a mistake you can decide to stop making. It operates on the reading, before any judgment you would recognize as a judgment, which is why noticing that an answer was well-argued is the moment to slow down rather than the moment to accept.
A. No, that pattern is not typical. Service accounts of this kind are normally scoped to server workloads, and interactive-style authentication to an end-user device would be unusual in most environments, though it can occur where legacy applications require it. I would treat it as worth reviewing.
B. I cannot tell. That depends on how your service accounts are scoped, which is not something I have.
B is correct and A is the one that reads like expertise. A hedges in the right places, names a legitimate exception, and closes with a recommendation, and every one of those properties is a property of the writing rather than of the claim.
This matters specifically because fluency is increasing. Each generation of these systems produces better-organized, better-hedged, more professional-sounding output, and none of that improvement is in the direction of being right about your estate. The signal that feels like reliability is getting stronger while its correlation with reliability stays where it was.
The rising line is the one you feel. The flat one is the one that decides whether the answer is right about your estate.
The practical form of this is worth recognizing in yourself. If your reason for accepting an answer is that it was well-argued, you have been persuaded rather than convinced, and the distinction is that a persuasive answer changes your belief without providing evidence.
The five conditions against a real request
Scoring a request that reads as completely reasonable, before any output existsTake a request an analyst made on the Northgate brute force and score it before reading the answer.
THE REQUEST, AS IT WAS TYPED
"The alert on r.scott fired around 23:00 on the 2nd. Show me what
happened with that account recently, and tell me if it looks
like the attack worked."
SCORE IT AGAINST THE FIVE BEFORE YOU READ ON.
Condition 1 fires twice. "Around 23:00" and "recently" are both relative, and the second has no anchor at all. Any window that misses 22:14 to 22:55 excludes the attack and still returns ordinary sign-ins.
Condition 2 fires. "Tell me if it looks like the attack worked" is a judgment about this estate. The answer will be about what successful brute forces generally look like.
Condition 3 is live. "The attack does not appear to have succeeded" is a negative claim, and it means the query it wrote surfaced no success rather than that none occurred.
Condition 4 is absent. No correlation across tables is required.
Condition 5 is likely. Any summary of "what happened with that account" will contain counts, and those counts will be generated.
SCORED, BEFORE ANY OUTPUT EXISTS
1 relative time FIRES TWICE "around 23:00", "recently"
2 about your estate FIRES "does it look like it worked"
3 answer could be LIVE a "no evidence" answer is a
a negative claim about the query, not
about the account
4 two sources absent no correlation needed
5 a figure to repeat LIKELY any summary carries counts
FOUR OF FIVE. NOTHING ABOUT THE REQUEST IS BADLY PHRASED.
Four of five, on a request that reads as completely reasonable. That is the point of the list: this is not a badly phrased request, it is how people actually ask, and the conditions fire on ordinary language rather than on carelessness.
The rewritten version costs nothing.
THE SAME QUESTION, ASKED SO THE CONDITIONS DO NOT FIRE
"Show me every authentication for r.scott@ne.com between
2026-03-02T18:00Z and 2026-03-03T06:00Z, including result
codes, sorted by time."
ABSOLUTE WINDOW. NO JUDGMENT REQUESTED. NO FIGURE TO GENERATE.
CONDITIONS 1, 2 AND 5 ARE GONE RATHER THAN CHECKED.
It produces the evidence the original question wanted and took five seconds longer to type.
Scoring the request rather than the answer moves the check earlier, where it is cheaper. Analysts who do this well are not more suspicious than everyone else. They are suspicious at a different point.
Recalibrating when the tool changes
Every calibration in this sub is about a specific tool at a specific time. When the model behind it changes, and it will without announcement, your hit rate and your sense of which requests are risky are both measured against a system that no longer exists.
WHAT AN UPGRADE CHANGES, AND WHAT IT DOES NOT
the six failure modes unchanged, they follow from
the mechanism
which of the six you meet CHANGES, and without notice
your hit rate measured against a system that
no longer exists
how well-written it reads IMPROVES, which is the trap
The failure modes do not change, because they follow from the mechanism. The distribution across them does. A new model might be better at time windows and no better at negative claims, which moves where your attention belongs without changing anything you would notice casually.
The cheap response is to re-run the twenty-request evaluation from the previous sub after any significant change, and to treat a sudden improvement in fluency as a reason to check more rather than less. An answer that got noticeably better-written did not necessarily get more correct, and the fluency trap is strongest immediately after an upgrade.
Test it on your own assistant
Write me a KQL query against SigninLogs to show what happened with that account around the alert.
What this demonstrates. Something will come back with a specific time range in it, invented from nothing, and the query will run against real data and return real rows. That is the silent window at its source.
Now rewrite the request with an absolute range, between (datetime("2026-03-02T18:00:00Z") .. datetime("2026-03-03T06:00:00Z")), and notice that the failure is gone rather than caught. Five seconds of rewriting removed a condition that would otherwise have needed a check.
The pre-flight
Reading your own request back, and which conditions a rewrite removesThe five conditions are only worth anything if you can run them at speed, and they run against your own sentence rather than against anything that came back.
READ YOUR OWN REQUEST BACK
1. RELATIVE TIME? "around the alert", "recently", "lately"
-> silent window. REWRITE IT, do not check it.
2. ABOUT YOUR ESTATE? "is this normal here", "does this look
suspicious" -> you will get a general truth.
3. COULD THE ANSWER a zero, a "no evidence of", an absence
BE A NEGATIVE? -> confident absence. Prove it can return rows.
4. TWO SOURCES? anything correlating tables -> wrong join key.
5. WILL YOU REPEAT any count, duration or total you will put in
A FIGURE? a ticket -> invented precision.
CONDITIONS 1, 2 AND 5 ARE REMOVED BY REWRITING THE REQUEST.
3 AND 4 ARE NOT, AND THOSE ARE THE ONES THAT EARN A CHECK.
Copy that somewhere you will see it. The checks you run on the output are in 0.3, one per failure mode. This card is the half that runs before anything has been generated.
The obvious objection is that plenty of requests are not yours. The card runs backwards: reconstruct the request that would have produced the artifact, then score that. A summary asserting what happened "recently" came from a request containing a relative time, whether or not you can see the sentence.
The three phrases that cost the most
Two of the five conditions are triggered by how the request was worded rather than by anything the model did. Three phrases account for most of it, and all three are how people naturally speak.
YOU SAY YOU GET SAY INSTEAD
─────────────────────────────────────────────────────────────────────
"around the alert" a window that may miss it "between 18:00Z
and 06:00Z"
"recently" / "lately" an anchor from today, not an absolute date
from the incident range
"is this suspicious?" a true statement about "show me every X
accounts in general for this account"
Rewriting the request is cheaper than checking the answer, and it removes the failure rather than catching it.
A working heuristic
Green, amber and red, and what each earnsEverything above compresses into something usable at speed.
Green. Translation, no relative time, no join, no figure to repeat. Read it against your intent and use it. This is most requests and it should feel fast.
Amber. Any of the five conditions present. Run the specific check that condition implies, not a general review.
Red is not a level of care. It is an instruction to go and get the evidence yourself rather than to check somebody else's.
Red. The request was about your estate, or the answer is going into a report, or the decision that follows is expensive to reverse. Do not verify the artifact. Go and get the evidence yourself, and use the generated answer as a hypothesis to test rather than a result to check.
Most artifacts are green. A habit that treats everything as amber costs more than it returns and gets dropped.
The bands are a way of deciding how much attention an artifact earns, not how serious the alert is. Getting that distinction right is what keeps the habit affordable, because most of what you see is green and treating it as amber is how the practice dies.
The Reasonable Mistake
The category people get wrong
Red is not "the important alerts". It is defined by whether the answer depends on your estate or is expensive to reverse, and plenty of routine alerts qualify while plenty of dramatic ones do not. An analyst who reserves care for incidents that feel serious will apply it to the wrong ones, because the alert that turns out to matter rarely announces itself in advance.
So the band is read off the request and the consequence, never off how the alert feels. Both of those you can name in a sentence before you have looked at anything, which is what makes the classification cheap enough to actually perform.
Why predicting beats checking
Where the cost sits, and why one of the two habits survives a shiftPrediction and checking are not two names for the same habit, and the difference is where the cost sits.
Checking happens after an artifact exists. It costs thirty seconds each time, it competes with everything else on your queue, and it is the step that goes first under pressure, because the reward for doing it is almost always nothing.
Predicting happens before, from your own sentence, and it costs about five seconds. It also changes what you do next rather than only what you conclude: a request exposed to a silent window is a request you rewrite with an absolute window, and the failure never occurs.
That is also why the five conditions are properties of the REQUEST rather than of the output. Anything you can only assess after reading the answer arrives too late to change the question, and by then you are in the expensive half of the work.
CHECKING PREDICTING
runs after the artifact before the request
exists is sent
costs ~30 seconds ~5 seconds
finds nothing, most nothing, most times
times
outcome you caught it it never happened
what goes first this one, because nothing. It is part of
under pressure the reward is typing the request
almost always
nothing
The bottom row is the one that decides it. A habit that competes with your queue loses to your queue, and the only version of this that survives a working month is the version that is part of writing the request rather than a step you perform afterwards.
Practice
Rewrite three requests before sending them hands onYou cannot calibrate against somebody else's hit rate. Five artifacts of your own, scored before you check them, is the smallest sample that tells you which half of the problem you actually have.
- Catch yourself about to type "around the alert" and give an absolute time range instead.
- Catch "recently" and give an absolute date range.
- Catch "is this normal?" and ask "how many times in ninety days?" instead.
Count how many times in one shift you were about to use one of the three. The number is usually higher than people expect, and it is the cheapest habit in this module because it costs five seconds and happens before anything has gone wrong.
Next: section 0.6 deals with the one risk that has nothing to do with correctness, which is what leaves your estate when you ask for help.