Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
Why Fluent AI Output Defeats an Analyst's Verification Instincts
Introduction
You have a working instinct for whether to trust an answer, and it is built on signals a person gives off: how long they took, whether they hedged, whether they volunteered how they know. This section is about what happens to that instinct when the answer arrives in four seconds, in the same confident register every time, from something with no track record and nothing at stake.
The short version is that your instinct does not fail, it fires on nothing. What replaces it has to be deliberate, and the reason it has to be deliberate is a specific fact about which comparison each kind of checking performs. By the end you will know why a careful read misses these failures every time, and what a check has to name to count as one.
Scenario
You review the query from 0.1 properly this time. You read every clause. The column names are real, the syntax is valid, the operations run in a sensible order, nothing is malformed. You approve it. It is wrong in exactly the way it was wrong before you looked, and your review changed nothing. This section is about why that happens to careful people rather than to careless ones.
The cues you lost
What a colleague gives off that generated text does notConsider how you evaluate information from a colleague. You do it constantly and almost none of it is conscious.
A junior analyst tells you the account was compromised at 02:14. You do not audit that claim from first principles; you read a dozen signals in parallel and mostly without noticing.
WHAT YOU READ FROM A COLLEAGUE, IN ABOUT A SECOND
how long they took to answer
whether they said "I think" or "it was"
whether they volunteered how they know
whether they are someone who overstates
whether this is the kind of claim that is easy to get wrong
what it would cost them to be wrong in front of you
NONE of these is about the claim. ALL of them arrive free.
From a generated answer, every one reads the same every time.
That system is efficient and it is almost entirely built on signals that have nothing to do with the claim itself. It works because human confidence and accuracy are correlated. Imperfectly, and enough to be useful.
Generated output arrives with every one of those signals stripped out.
EVERY SIGNAL, AND WHAT IT READS AS FROM GENERATED TEXT
latency hard question and trivial one: both 4 seconds
register fluency is a property of the text, not of any
assessment of the claim
volunteered only when asked, and then it sounds equally
reasoning plausible whether or not it is the reasoning
that produced the answer
track record none. No reputation to protect either.
Why this is not a diligence problem. The instinct that fails you here is a good instinct. Reading confidence as a proxy for reliability is a sound heuristic among humans and it is what makes working with colleagues possible at all. It fails with generated output for a specific structural reason: the correlation it depends on has been removed. Blaming the analyst for applying it is like blaming someone for trusting a clock that turns out to be stopped.
Five signals on the left, none of them present on the right. Your instinct did not fail here, it had nothing to read.
Why checking feels done when it is not
Reading tests the artifact against itself query to runThere is a second effect, and it is more insidious than the first because it operates after you have decided to be careful.
When you review a generated query, you read it. It parses. The column names are real. The logic is coherent: a filter, a summarize, a sort, in a sensible order. Nothing is obviously absurd. You have now performed a review, and it felt like one, because you did in fact examine the artifact attentively for several seconds.
What you checked was plausibility. What matters is correspondence: whether this query, run against this data, answers the question you actually asked. Those are different properties and only one of them is visible on the page.
Plausibility is cheap to assess and it is what the eye does automatically. Correspondence requires you to hold your original question in mind, work out what the query would return, and compare. That is real cognitive work, it takes longer than reading, and there is nothing in the experience of reading a well-formed query that tells you that you have not done it.
The Reasonable Mistake
The specific trap
A generated artifact that is badly formed gets checked properly, because the mess prompts you to slow down. A generated artifact that is beautifully formed gets a nod. The output most likely to be accepted without correspondence checking is the output that looks best, and there is no relationship between how good a query looks and whether it answers your question.
The example on the orientation page is exactly this. Nothing about where ResultType == 0 looks wrong. It is a real column, a valid comparison, and filtering by result is precisely what a thoughtful analyst would do. It reads as diligence. You have to hold the original request in mind, remember that the analyst asked about failures, and notice that zero means success, before it becomes visible.
See it happen, on one query
Reading about plausibility is not the same as being caught by it. Here is a query. Give it ten seconds, decide whether you would run its result into a ticket, then read on.
SigninLogs
| where ConditionalAccessStatus == "notApplied"
| where IPAddress !startswith "10.0."
| summarize AffectedUsers = count() by UserPrincipalName, AppDisplayName
Most people accept it. It is clean, the column names are real, the intent is legible, and the variable is honestly named. Run it and two rows come back, both service accounts using Azure AD PowerShell from outside the estate. That is a genuine finding and the query found it.
Now read it again with your original question in mind. AffectedUsers = count() counts EVENTS and labels them users. With two rows and small numbers nothing looks wrong. Ask the same query about a wider filter and the label starts reporting a number that is not the thing it claims to be.
The Reasonable Mistake
Checking that a query is well-formed instead of checking what it counts
Across the whole estate, that pattern returns 12 events from 11 distinct users. The gap is one, which is exactly the size of gap that never gets questioned and always survives into a report.
The correct form is dcount(UserPrincipalName). The difference is one function call, invisible on the page, and it is the difference between a count of people and a count of log lines.
That is plausibility checking failing in real time. You read it, it looked right, and it was well-formed and mislabelled at once. Nothing about reading it more carefully would have helped, because the error is in the relationship between the label and the function, not in either one.
The mechanism: what reading tests, and what it does not
The reason is not effort. It is which comparison each kind of checking performs.
Both sides of the left comparison sit on the screen, which is why it never fails and never catches anything.
That asymmetry has a second consequence, and it is the one that decides whether a check is worth building.
Every one of the six failure modes is invisible to the first comparison and visible to the second, and that is not a coincidence. They come from a process that generates well-formed text, so a check testing well-formedness is testing the one property guaranteed to pass.
WHAT EACH COMPARISON WOULD HAVE CAUGHT
mode reading it checking it
──────────────────────────────────────────────────────
plausible field no yes
silent window no yes
wrong join key no yes
confident absence no yes
right answer, wrong question no yes
invented precision no yes
malformed syntax YES yes
a column that does not exist YES yes
Why this gets worse as the tooling improves
The gap between the two comparisons widens as generated output gets better, which is the opposite of what most people assume.
When a tool produced obviously broken output, reading was an adequate check: the mess was on the page and the first comparison caught it. As the output became well-formed nearly all of the time, the first comparison stopped catching anything, while continuing to feel exactly as reassuring as it always had.
The signal did not weaken. The thing it was tracking changed underneath it. An analyst whose habit was formed on earlier tooling is running a check calibrated for a failure mode that no longer occurs, against a failure mode it was never able to see.
The cost asymmetry
Skipping is rewarded now, and the bill arrives lateThe economics explain why this persists in teams that know about it.
Verification is paid every time. The error is paid rarely.
THE ECONOMICS OF SKIPPING IT
cost how often
──────────────────────────────────────────────────────────
checking an artifact 90 seconds EVERY time
30 artifacts a shift 45 minutes every shift
finds nothing most of them because most output is fine
a wrong query on a nothing you never find out
benign alert
a wrong query on the a missed rarely, and it is the
one that mattered intrusion one that counts
Why the feedback loop cannot teach you this. Skipping verification is rewarded immediately and reliably: you save the ninety seconds and nothing bad happens. Doing it is punished immediately and mildly: you spend the time and usually find nothing. The consequence of skipping arrives late, rarely, and disconnected from the decision that caused it, so nothing in your day-to-day experience will ever teach you the habit. This is also why "be careful with AI output" is useless advice. Everyone agrees with it and the incentive structure defeats it within a week. What survives is a check specific enough to be quick.
The three states a reviewer can be in
It helps to be able to name where you are, because two of these feel identical from the inside.
The middle state is the dangerous one, because from the inside it is indistinguishable from the third.
What this means for a handover note. Write what you checked, not that you checked. "Reviewed the query" is the second state wearing the language of the third. "Confirmed the filter selects failures rather than successes" is checkable by whoever reads it, and it tells them exactly which risk you have retired and, by omission, which you have not.
INSTEAD OF WRITE
─────────────────────────────────────────────────────────────
Reviewed the query Confirmed ResultType 50126 is
a failure code, not a success
Looks correct Ran it with the window widened
a day each way. Same shape.
Query returned nothing Dropped the account filter and
got 41 rows, so the table and
window are right and the zero
is real
Checked the summary The 14-file figure is not in the
source data. Unverified.
Why more careful reading does not fix it
The obvious response to all of this is to read the output more carefully, and it does not work. Two reasons, both structural.
Careful reading finds a different class of error. Attention catches a mismatched bracket or a column that does not exist. It cannot catch an error that lives in the relationship between the artifact and the schema, the data, or your original question, because none of those three is on the page.
Attention does not survive volume. Whatever your review standard is at the first alert of a shift, it is lower at the twenty-fifth, and the twenty-fifth is where the intrusion is. A check that depends on sustained care degrades exactly when it is most needed, which is the same reason runbooks exist for incident response rather than trusting experienced responders to remember.
A finishable check holds its level because finishing it does not depend on how much attention you have left.
What survives volume is a short, specific, finishable check. Not "read it carefully" but "confirm the filter matches the intent, confirm the window covers the event, confirm an empty result is really empty". Three questions with answers, which is why the next sub is a list of six rather than an exhortation to be thorough.
The cost that arrives late
Nothing in a shift connects the two columns, which is why experience never teaches this one.
Nothing in a shift will ever teach you this habit, which is why it has to be installed deliberately.
Test it on your own assistant
service accounts logging on to workstations rather than servers.
Then tell me how confident you are that this query is correct, and why.
What this demonstrates. The confidence will read as considered and will usually be high. It is generated by the same process that generated the query, from the same probabilities, which is why it cannot function as a second opinion on it. A colleague's hesitation is evidence because it costs them something to admit; this costs nothing.
The query itself is worth a look too. Distinguishing a workstation from a server is estate knowledge, and nothing in the request supplied it, so whatever rule it used to tell them apart was invented.
What replaces the missing cue
A named check, not a resolution to be carefulYou cannot restore the signals you lost. Prompting for uncertainty produces text about uncertainty rather than a measure of it.
SIX QUESTIONS, NOT "REVIEW IT CAREFULLY"
1 does the variable name agree with the filter above it
2 is the window absolute, anchored on an EVENT time
3 is this the table where that behavior is recorded here
4 if it returned nothing, can it return rows at all
5 is any number counting what the question asked
6 did this answer the question, or one next to it
Bounded. Answerable. Finished in ninety seconds.
"Review this carefully" is none of those.
What you can do is stop relying on a signal at all, and instead ask a question whose answer is checkable.
The reason this works is that generated errors are not random. They cluster into a small number of shapes, because they arise from the same mechanism every time: the model produced the likely continuation rather than the true one. Likely-but-untrue has a limited vocabulary in a query language. It looks like the common column rather than the correct one, the common time field, the common join, the shape of an answer to a nearby question.
That is why the next sub can name six failure modes and claim the list is close to exhaustive for practical purposes. Once you know the shapes, checking stops being "review this carefully", which is unbounded and therefore never finished, and becomes six specific questions you can answer in ninety seconds.
READING tests the artifact against ITSELF
does the syntax parse
do the clauses agree with each other
does the description match the logic
-> a generated artifact passes all three by construction
CHECKING tests the artifact against SOMETHING ELSE
against the question you asked
against what the data actually contains
against a fact about your estate
-> this is the comparison that can fail
THE TELL if your check did not name something outside the
artifact, you were reading.
The tell at the foot of that card is the part worth carrying. It is a question about your own process rather than about the artifact, and it has a definite answer: name the thing you compared the artifact against. If you cannot, you read it.
The one question
What would have to be true for this to be rightUnderneath all six is a single question, and if you take one thing from this module it should be this one:
Not "does this look right", which is plausibility. Not "did it run", which is syntax. The question forces you to name a claim about the world and then check it.
Applied to the orientation example, it produces: for this to answer my question, ResultType == 0 would have to select failures. That is now a checkable statement, and thirty seconds with the schema settles it.
Applied to a detection rule: for this to be a useful rule, the behavior it matches would have to be rare in my estate. Checkable, by running it against a week of data.
Applied to a summary: for this summary to be complete, nothing outside it would have to change the finding. Checkable, by looking at what it left out.
The failed sign-ins for this account originate from a single ISP range in Vilnius and the pattern is consistent with a password spray rather than targeted credential stuffing.
Two claims, resolved differently. For the first, every address would have to belong to one range, which is one summarize away. For the second, one source would have to be trying a few passwords across many accounts, which this query never looked at. The second claim is not wrong. It is unsupported, and those need different responses.
Sometimes the statement is not checkable in the time you have. Say so: an artifact whose condition you cannot test gets escalated or declined, and naming which beats a review that ends in a shrug.
The question is the same every time. What changes is what you check, and that is what the six failure modes give you.
Practice
Measure your own hit rate over five artifacts hands onKnowing the difference between the two comparisons changes nothing on its own. Measuring your own hit rate over five artifacts is what tells you which half of the problem you actually have.
- For the next five generated queries you review, write one word before you check: sound, or unsound.
- Then check each one properly.
- Count how many you called correctly.
That number is your calibration, and it is worth more than a general resolution to be careful. If you catch four in five on identity queries and two in five on endpoint ones, you have learned exactly where your extra thirty seconds belongs.
Next: section 0.3 names the six shapes these failures take, each with the check that settles it, so "review it carefully" becomes something with an end.