In this section

Why Fluent AI Output Defeats an Analyst's Verification Instincts

Module 0

Introduction

You have a working instinct for whether to trust an answer, and it is built on signals a person gives off: how long they took, whether they hedged, whether they volunteered how they know. This section is about what happens to that instinct when the answer arrives in four seconds, in the same confident register every time, from something with no track record and nothing at stake.

The short version is that your instinct does not fail, it fires on nothing. What replaces it has to be deliberate, and the reason it has to be deliberate is a specific fact about which comparison each kind of checking performs. By the end you will know why a careful read misses these failures every time, and what a check has to name to count as one.

Scenario

You review the query from 0.1 properly this time. You read every clause. The column names are real, the syntax is valid, the operations run in a sensible order, nothing is malformed. You approve it. It is wrong in exactly the way it was wrong before you looked, and your review changed nothing. This section is about why that happens to careful people rather than to careless ones.

01

The cues you lost

What a colleague gives off that generated text does not

Consider how you evaluate information from a colleague. You do it constantly and almost none of it is conscious.

A junior analyst tells you the account was compromised at 02:14. You do not audit that claim from first principles; you read a dozen signals in parallel and mostly without noticing.

WHAT YOU READ FROM A COLLEAGUE, IN ABOUT A SECOND
 
  how long they took to answer
  whether they said "I think" or "it was"
  whether they volunteered how they know
  whether they are someone who overstates
  whether this is the kind of claim that is easy to get wrong
  what it would cost them to be wrong in front of you
 
NONE of these is about the claim. ALL of them arrive free.
 
From a generated answer, every one reads the same every time.

That system is efficient and it is almost entirely built on signals that have nothing to do with the claim itself. It works because human confidence and accuracy are correlated. Imperfectly, and enough to be useful.

Generated output arrives with every one of those signals stripped out.

EVERY SIGNAL, AND WHAT IT READS AS FROM GENERATED TEXT
 
latency          hard question and trivial one: both 4 seconds
register         fluency is a property of the text, not of any
                 assessment of the claim
volunteered      only when asked, and then it sounds equally
reasoning        plausible whether or not it is the reasoning
                 that produced the answer
track record     none. No reputation to protect either.

Why this is not a diligence problem. The instinct that fails you here is a good instinct. Reading confidence as a proxy for reliability is a sound heuristic among humans and it is what makes working with colleagues possible at all. It fails with generated output for a specific structural reason: the correlation it depends on has been removed. Blaming the analyst for applying it is like blaming someone for trusting a clock that turns out to be stopped.

THE SIGNALS YOU LOSE A COLLEAGUE Latency varies with difficulty Hedges when unsure Volunteers how they know Has a track record with you Has a reputation to protect GENERATED OUTPUT Four seconds, always Same register, always Reasoning on request, plausible No history Nothing at stake

Five signals on the left, none of them present on the right. Your instinct did not fail here, it had nothing to read.

02

Why checking feels done when it is not

Reading tests the artifact against itself
query to run

There is a second effect, and it is more insidious than the first because it operates after you have decided to be careful.

When you review a generated query, you read it. It parses. The column names are real. The logic is coherent: a filter, a summarize, a sort, in a sensible order. Nothing is obviously absurd. You have now performed a review, and it felt like one, because you did in fact examine the artifact attentively for several seconds.

What you checked was plausibility. What matters is correspondence: whether this query, run against this data, answers the question you actually asked. Those are different properties and only one of them is visible on the page.

Reading tests the artifact against itself. Only one of those two properties can fail on the page, and it is not the one that matters.

Plausibility is cheap to assess and it is what the eye does automatically. Correspondence requires you to hold your original question in mind, work out what the query would return, and compare. That is real cognitive work, it takes longer than reading, and there is nothing in the experience of reading a well-formed query that tells you that you have not done it.

The Reasonable Mistake

The specific trap

A generated artifact that is badly formed gets checked properly, because the mess prompts you to slow down. A generated artifact that is beautifully formed gets a nod. The output most likely to be accepted without correspondence checking is the output that looks best, and there is no relationship between how good a query looks and whether it answers your question.

The example on the orientation page is exactly this. Nothing about where ResultType == 0 looks wrong. It is a real column, a valid comparison, and filtering by result is precisely what a thoughtful analyst would do. It reads as diligence. You have to hold the original request in mind, remember that the analyst asked about failures, and notice that zero means success, before it becomes visible.

See it happen, on one query

Reading about plausibility is not the same as being caught by it. Here is a query. Give it ten seconds, decide whether you would run its result into a ticket, then read on.

SigninLogs
| where ConditionalAccessStatus == "notApplied"
| where IPAddress !startswith "10.0."
| summarize AffectedUsers = count() by UserPrincipalName, AppDisplayName

Most people accept it. It is clean, the column names are real, the intent is legible, and the variable is honestly named. Run it and two rows come back, both service accounts using Azure AD PowerShell from outside the estate. That is a genuine finding and the query found it.

Now read it again with your original question in mind. AffectedUsers = count() counts EVENTS and labels them users. With two rows and small numbers nothing looks wrong. Ask the same query about a wider filter and the label starts reporting a number that is not the thing it claims to be.

The Reasonable Mistake

Checking that a query is well-formed instead of checking what it counts

Across the whole estate, that pattern returns 12 events from 11 distinct users. The gap is one, which is exactly the size of gap that never gets questioned and always survives into a report.

The correct form is dcount(UserPrincipalName). The difference is one function call, invisible on the page, and it is the difference between a count of people and a count of log lines.

That is plausibility checking failing in real time. You read it, it looked right, and it was well-formed and mislabelled at once. Nothing about reading it more carefully would have helped, because the error is in the relationship between the label and the function, not in either one.

The mechanism: what reading tests, and what it does not

The reason is not effort. It is which comparison each kind of checking performs.

WHICH COMPARISON IS EACH ONE MAKING? READING the artifact the artifact both sides on the screen. Fast, automatic, and it always passes. CHECKING the artifact your original question not on the screen neither side is visible. Deliberate, and the step skipped under load. A generated artifact is coherent by construction, so the left comparison cannot fail.

Both sides of the left comparison sit on the screen, which is why it never fails and never catches anything.

That asymmetry has a second consequence, and it is the one that decides whether a check is worth building.

Every one of the six failure modes is invisible to the first comparison and visible to the second, and that is not a coincidence. They come from a process that generates well-formed text, so a check testing well-formedness is testing the one property guaranteed to pass.

WHAT EACH COMPARISON WOULD HAVE CAUGHT
 
mode                        reading it   checking it
──────────────────────────────────────────────────────
plausible field             no           yes
silent window               no           yes
wrong join key              no           yes
confident absence           no           yes
right answer, wrong question no          yes
invented precision          no           yes
 
malformed syntax            YES          yes
a column that does not exist YES         yes

Why this gets worse as the tooling improves

The gap between the two comparisons widens as generated output gets better, which is the opposite of what most people assume.

When a tool produced obviously broken output, reading was an adequate check: the mess was on the page and the first comparison caught it. As the output became well-formed nearly all of the time, the first comparison stopped catching anything, while continuing to feel exactly as reassuring as it always had.

The signal did not weaken. The thing it was tracking changed underneath it. An analyst whose habit was formed on earlier tooling is running a check calibrated for a failure mode that no longer occurs, against a failure mode it was never able to see.

03

The cost asymmetry

Skipping is rewarded now, and the bill arrives late

The economics explain why this persists in teams that know about it.

Verification is paid every time. The error is paid rarely.

THE ECONOMICS OF SKIPPING IT
 
                        cost           how often
──────────────────────────────────────────────────────────
checking an artifact    90 seconds     EVERY time
30 artifacts a shift    45 minutes     every shift
finds nothing           most of them   because most output is fine
 
a wrong query on a      nothing        you never find out
benign alert
a wrong query on the    a missed       rarely, and it is the
one that mattered       intrusion      one that counts

Why the feedback loop cannot teach you this. Skipping verification is rewarded immediately and reliably: you save the ninety seconds and nothing bad happens. Doing it is punished immediately and mildly: you spend the time and usually find nothing. The consequence of skipping arrives late, rarely, and disconnected from the decision that caused it, so nothing in your day-to-day experience will ever teach you the habit. This is also why "be careful with AI output" is useless advice. Everyone agrees with it and the incentive structure defeats it within a week. What survives is a check specific enough to be quick.

The three states a reviewer can be in

It helps to be able to name where you are, because two of these feel identical from the inside.

TWO OF THESE FEEL IDENTICAL FROM THE INSIDE UNCHECKED read it, it looked fine, ran it you know you did not verify anything HONEST PLAUSIBILITY-CHECKED read attentively, found nothing malformed feels like verification. Is not. NOT HONEST, AND SINCERE CORRESPONDENCE-CHECKED named what would have to be true, confirmed it takes longer, and the answer is defensible HONEST The middle state is the one the six failure modes are built to survive. "I did not check it" is accurate. "I reviewed it" from the middle state sounds like the right-hand one, and the person saying it believes it.

The middle state is the dangerous one, because from the inside it is indistinguishable from the third.

What this means for a handover note. Write what you checked, not that you checked. "Reviewed the query" is the second state wearing the language of the third. "Confirmed the filter selects failures rather than successes" is checkable by whoever reads it, and it tells them exactly which risk you have retired and, by omission, which you have not.

Keep this Handover wording that says which state you were in
INSTEAD OF                     WRITE
─────────────────────────────────────────────────────────────
Reviewed the query             Confirmed ResultType 50126 is
                               a failure code, not a success
 
Looks correct                  Ran it with the window widened
                               a day each way. Same shape.
 
Query returned nothing         Dropped the account filter and
                               got 41 rows, so the table and
                               window are right and the zero
                               is real
 
Checked the summary            The 14-file figure is not in the
                               source data. Unverified.
Each right-hand entry names one risk you retired. Whatever it does not name is still open, and the next reader can see which.

Why more careful reading does not fix it

The obvious response to all of this is to read the output more carefully, and it does not work. Two reasons, both structural.

Careful reading finds a different class of error. Attention catches a mismatched bracket or a column that does not exist. It cannot catch an error that lives in the relationship between the artifact and the schema, the data, or your original question, because none of those three is on the page.

Attention does not survive volume. Whatever your review standard is at the first alert of a shift, it is lower at the twenty-fifth, and the twenty-fifth is where the intrusion is. A check that depends on sustained care degrades exactly when it is most needed, which is the same reason runbooks exist for incident response rather than trusting experienced responders to remember.

ONE OF THESE DEPENDS ON HOW YOU FEEL AT ALERT 25 high low alert 1 alert 12 alert 25 read it carefully unbounded, so it is the first thing to give three named checks, ninety seconds the intrusion is at alert 25

A finishable check holds its level because finishing it does not depend on how much attention you have left.

What survives volume is a short, specific, finishable check. Not "read it carefully" but "confirm the filter matches the intent, confirm the window covers the event, confirm an empty result is really empty". Three questions with answers, which is why the next sub is a list of six rather than an exhortation to be thorough.

The cost that arrives late

WHY THE FEEDBACK LOOP TEACHES THE WRONG LESSON CHECKING 90 seconds, 30 times a shift, and almost always finds nothing -45m NOT CHECKING Saves the 90 seconds. Nothing happens. Nothing happens. Nothing happens. +45m then once the alert mattered Skipping is rewarded immediately and reliably. The cost lands late, rarely, and looks like the alert's fault.

Nothing in a shift connects the two columns, which is why experience never teaches this one.

Nothing in a shift will ever teach you this habit, which is why it has to be installed deliberately.

Test it on your own assistant

Try this Ask it how confident it is, then ask again
Write me a KQL query against a table called DeviceLogonEvents that finds

service accounts logging on to workstations rather than servers.

Then tell me how confident you are that this query is correct, and why.

What to look at. The confidence statement, and whether it changes anything. Ask the same question a second time in a new conversation and compare the two. A stated confidence that does not vary with the difficulty of the question is not information about the answer.

What this demonstrates. The confidence will read as considered and will usually be high. It is generated by the same process that generated the query, from the same probabilities, which is why it cannot function as a second opinion on it. A colleague's hesitation is evidence because it costs them something to admit; this costs nothing.

The query itself is worth a look too. Distinguishing a workstation from a server is estate knowledge, and nothing in the request supplied it, so whatever rule it used to tell them apart was invented.

04

What replaces the missing cue

A named check, not a resolution to be careful

You cannot restore the signals you lost. Prompting for uncertainty produces text about uncertainty rather than a measure of it.

SIX QUESTIONS, NOT "REVIEW IT CAREFULLY"
 
  1  does the variable name agree with the filter above it
  2  is the window absolute, anchored on an EVENT time
  3  is this the table where that behavior is recorded here
  4  if it returned nothing, can it return rows at all
  5  is any number counting what the question asked
  6  did this answer the question, or one next to it
 
Bounded. Answerable. Finished in ninety seconds.
"Review this carefully" is none of those.

What you can do is stop relying on a signal at all, and instead ask a question whose answer is checkable.

The reason this works is that generated errors are not random. They cluster into a small number of shapes, because they arise from the same mechanism every time: the model produced the likely continuation rather than the true one. Likely-but-untrue has a limited vocabulary in a query language. It looks like the common column rather than the correct one, the common time field, the common join, the shape of an answer to a nearby question.

That is why the next sub can name six failure modes and claim the list is close to exhaustive for practical purposes. Once you know the shapes, checking stops being "review this carefully", which is unbounded and therefore never finished, and becomes six specific questions you can answer in ninety seconds.

Reference Reading against checking the whole difference
READING tests the artifact against ITSELF
  does the syntax parse
  do the clauses agree with each other
  does the description match the logic
  -> a generated artifact passes all three by construction
 
CHECKING tests the artifact against SOMETHING ELSE
  against the question you asked
  against what the data actually contains
  against a fact about your estate
  -> this is the comparison that can fail
 
THE TELL   if your check did not name something outside the
           artifact, you were reading.
Every one of the six failure modes survives the first column and dies in the second.

The tell at the foot of that card is the part worth carrying. It is a question about your own process rather than about the artifact, and it has a definite answer: name the thing you compared the artifact against. If you cannot, you read it.

05

The one question

What would have to be true for this to be right

Underneath all six is a single question, and if you take one thing from this module it should be this one:

What would have to be true for this answer to be correct?

Not "does this look right", which is plausibility. Not "did it run", which is syntax. The question forces you to name a claim about the world and then check it.

Applied to the orientation example, it produces: for this to answer my question, ResultType == 0 would have to select failures. That is now a checkable statement, and thirty seconds with the schema settles it.

Applied to a detection rule: for this to be a useful rule, the behavior it matches would have to be rare in my estate. Checkable, by running it against a week of data.

Applied to a summary: for this summary to be complete, nothing outside it would have to change the finding. Checkable, by looking at what it left out.

Generated triage note · assistant output

The failed sign-ins for this account originate from a single ISP range in Vilnius and the pattern is consistent with a password spray rather than targeted credential stuffing.

Two claims, resolved differently. For the first, every address would have to belong to one range, which is one summarize away. For the second, one source would have to be trying a few passwords across many accounts, which this query never looked at. The second claim is not wrong. It is unsupported, and those need different responses.

Sometimes the statement is not checkable in the time you have. Say so: an artifact whose condition you cannot test gets escalated or declined, and naming which beats a review that ends in a shrug.

The question is the same every time. What changes is what you check, and that is what the six failure modes give you.

06

Practice

Measure your own hit rate over five artifacts
hands on

Knowing the difference between the two comparisons changes nothing on its own. Measuring your own hit rate over five artifacts is what tells you which half of the problem you actually have.

Practice Measure your own hit rate
  1. For the next five generated queries you review, write one word before you check: sound, or unsound.
  2. Then check each one properly.
  3. Count how many you called correctly.
Most people score lower than they expect on the artifacts that look fine, because the cases they remember are the ones they caught.

That number is your calibration, and it is worth more than a general resolution to be careful. If you catch four in five on identity queries and two in five on endpoint ones, you have learned exactly where your extra thirty seconds belongs.

Next: section 0.3 names the six shapes these failures take, each with the check that settles it, so "review it carefully" becomes something with an end.