In this section

Why AI-Generated Security Queries Fail: The Mechanism

Module 0

Introduction

An AI assistant will write you a query in four seconds that would have taken you fifteen minutes, and most of the time it will be right. This section is about the mechanism that produces the rest: not a list of mistakes to memorize, but the single behavior underneath all of them, which is that these systems emit the most probable continuation rather than the correct one.

That distinction sounds academic and it is the most practical thing in the module. Understand it and you can look at a query against a table this course never mentions, or a tool that does not exist yet, and say which clause is most likely wrong before you run anything. By the end you will have that rule, tested against three different fields in the Northgate estate.

Scenario

An alert fires on r.scott@ne.com on the night of 2 March. You ask an assistant for the failed sign-ins on that account so you can see whether the attack succeeded. The query comes back in four seconds, runs without error, and returns rows. You need to decide whether to act on it, and the only thing in front of you is the query itself.

01

What the model is actually doing

Probable continuation, and what "probable" is measured against
query to run

A language model produces text one token at a time, and each token is the most probable continuation given everything before it. That is the entire mechanism, and every failure in this course falls out of it.

"Probable" means probable in what the model was trained on, which for a query language is documentation, tutorials, Stack Overflow answers, vendor samples and public repositories. It does not mean probable given your estate, your alert, or the sentence you just typed, except insofar as your sentence shifts the odds.

That last clause is where the trouble is. Your request influences the output, and it competes against everything the model has ever seen. When your wording and the training distribution disagree, the distribution usually wins, because it is thousands of examples against one sentence.

Watch it happen on the scenario above.

SigninLogs
| where UserPrincipalName == "r.scott@ne.com"
| where ResultType == 0
| summarize Failures = count() by IPAddress
ONE TOKEN POSITION, TWO FORCES YOUR REQUEST "the failed sign-ins" weight: one sentence TRAINING DISTRIBUTION ResultType == 0 in every sample weight: thousands next token after where ResultType == 0 The distribution wins because it is thousands of examples against one sentence.

Only one token position in this query is contested. The other clauses had nothing competing with what you asked for.

Run it. Ten rows come back, all office addresses, 130 events in total.

You asked for failures. ResultType == 0 is success. The query counts successes and labels them Failures.

02

Why that error, and not a random one

Where your wording loses, and to what

The error is not arbitrary, and understanding why makes it predictable.

ResultType is the right column: it is the field that records sign-in outcome, and the model has seen it in that role thousands of times. The failure is in the value. In Microsoft's documentation, in every example query, in every tutorial that shows a successful sign-in, the comparison overwhelmingly written is ResultType == 0. The value zero is the default case, the happy path, the one that appears in the sample everybody copies.

Your word "failed" appeared once, in your request.

One sentence of yours against a large body of examples, on a single token. The model resolves that the way it resolves everything, by probability.

Now look at what did carry your intent. The variable is called Failures, generated by the same process moments later, and it kept your word because nothing competed for that position.

WHY ONE CLAUSE KEPT YOUR WORD AND THE OTHER DID NOT | where ResultType == 0 | summarize Failures = count() by IPAddress A CONVENTION EXISTS HERE Result-code comparisons appear thousands of times, nearly always as == 0. Your word loses. NO CONVENTION HERE Variable names are arbitrary, so nothing competes with what you asked for. Your word wins.

One query, two clauses, two outcomes. A convention can only beat your wording where a convention exists.

The Reasonable Mistake

Reading the label as evidence of the logic

The variable name and the filter were produced by the same process moments apart, so it feels safe to assume they agree. They frequently do not, and the disagreement is systematic rather than random: names follow your request because names have no convention, and filters follow the training distribution because filters do. A mismatch between a name and the clause above it is one of the highest-yield signals available to you, and it exists because of how the two positions differ.

So the two positions disagree for a reason you can state, and the disagreement is visible without running anything: read the name, read the filter above it, and see whether they describe the same thing.

03

Where your wording does win

The positions with no convention to compete against, and why they need no check

The picture so far makes the model sound like it ignores you. It does not, and knowing where your request reliably wins is as useful as knowing where it loses, because it tells you which parts of a query you can stop checking.

Your wording dominates wherever there is no convention to compete with. Nothing in the training data has an opinion about whether you meant r.scott or k.foster, so the token that appears is the one you supplied. The same holds for the projection and the sort order, which are properties of your question rather than of the language.

The Reasonable Mistake

Spreading the check evenly across the query

The careful reviewer reads every clause with equal attention, which sounds like diligence and is the reason the contested one gets the same four seconds as the account name you typed yourself. Attention is finite and the clauses are not equally at risk, so even attention means the position most likely to be wrong is the position you looked at least hard relative to its odds.

Structure is a middle case. Whether the query summarizes or projects, whether it filters before aggregating, follows the shape of your request fairly reliably, because you described a shape and shapes are what the model is matching. The overall form came from you; the individual values came from the distribution, which is why a well-described query usually has the right skeleton and the wrong detail rather than the reverse.

That has a measurable consequence for how long a review takes. A ten-clause query typically has two or three contested positions and seven that carried your intent, so reading all ten costs you the same minute and finds less. Reading three costs twenty seconds and covers the same risk.

WHERE TO SPEND YOUR THIRTY SECONDS SigninLogs | where UserPrincipalName == "r.scott@ne.com" yours, safe | where ResultType == 0 CONTESTED | summarize Failures = count() by IPAddress yours, safe | sort by Failures desc yours, safe CHECK THESE Result codes, status strings, time columns, table names, join keys LEAVE THESE ALONE Literals you supplied, projection, sort order, overall structure

Three of these four clauses carried your intent unchanged. Spend the thirty seconds on the one that did not.

04

Why more context helps, and where it stops

What a standing preamble buys you, and the facts it cannot reach

The obvious response is to tell the model more: give it the schema, describe the estate, spell out that zero means success in this table. That helps, and it helps in a way the mechanism predicts. Context does not remove the training distribution, it adds weight to your side of the same scale, which is why it helps most where the convention is weak and least where it is overwhelming. It never reaches a fact that was never in the model: nothing you write tells it that svc-legacyerp has authenticated nightly for two years without anybody minding.

A short standing preamble is worth writing once. Ten lines describing your tables, your ranges and the two or three fields whose values are counter-intuitive, pasted at the start of a session.

Keep this A standing preamble, written once and pasted every time
TABLES I USE
  SigninLogs                        interactive sign-ins
  AADNonInteractiveUserSignInLogs   token-backed, where replay lives
  DeviceProcessEvents               process creation, endpoints
 
FIELDS WHOSE VALUES READ BACKWARDS
  ResultType        0 is SUCCESS. Failures are non-zero.
  ConditionalAccessStatus  notApplied is not a failure
 
ADDRESS RANGES
  10.0.x.x   corporate. Anything else is external.
 
WHEN I ASK FOR FAILURES I MEAN FAILURES, NOT THE DEFAULT CASE.
Yours will name different fields. The two or three whose values read backwards are the ones worth the line.

Writing it once changes the odds on every request you make afterwards, which is why it is worth the ten minutes. What it does not do is remove the contested positions, so the check still runs. It runs against fewer clauses.

05

The rule this gives you

Expect the documented default, whatever you asked for
query to run

Everything above compresses into something you can apply to a table you have never opened.

Where a field has a dominant value in the documentation, expect the dominant value, whatever you asked for.

That is testable. Try it on a different field, in a different table, on the same estate.

ConditionalAccessStatus has three possible values: success, failure, and notApplied. Which one dominates the documentation? success, overwhelmingly, because that is what a working policy produces and what every example shows. So a request for sign-ins where policy did not apply is exactly the shape that goes wrong.

SigninLogs
| where ConditionalAccessStatus == "notApplied"
| summarize Gaps = count()

This one is correct: 12 events. But you can now say in advance that it was a coin-flip clause, and that the check worth thirty seconds is confirming the status value rather than reading the number. You knew which clause to check before you ran it, and you knew it from the shape of the field rather than from experience with this query.

Test it on your own assistant

Everything above is a claim about how these systems behave. Here is how to check it yourself, in about a minute, with whatever assistant you already use.

Try this Ask for failures and see which value comes back
I have a Microsoft Entra ID sign-in log table called SigninLogs, with columns

UserPrincipalName, IPAddress, ResultType and TimeGenerated.

Write me a KQL query that counts the FAILED sign-ins for one user, grouped by IP address.

Nothing in that request identifies anyone. No account, no hostname, no address, which is deliberate: 0.6 covers why estate identifiers should not leave, and an activity that told you to paste yours would contradict the module.

What to look at. One thing only: the comparison on ResultType. You asked for failures. A result code of zero is success, so a correct query needs != 0. Check which one you got, and check whether the variable it chose agrees with the filter above it.

What usually comes back. Most of the time the query is correct, because the word FAILED in capitals is doing a lot of work. Weaken the request to "show me the failed sign-ins" and the odds shift. That sensitivity is the mechanism in this section: your wording competes with the documented default, and it wins or loses at the margin rather than reliably.

YOUR WORDING COMPETES WITH THE DOCUMENTED DEFAULT "count the FAILED sign-ins" usually != 0 your word wins "show me the failed sign-ins" it varies the margin "summarize sign-in results" the default wins: ResultType == 0 The clause at risk is the one where a convention competes with your request.

The same request in three registers. Emphasis moves the margin without removing the competition.

A correct result is not a refutation. You have just watched which clause was contested in a case where nothing failed, and that is the transferable part.

06

The reach of the mechanism

Why the list is six long rather than sixty, and the two failures it does not cover

The dominant-value effect is one form of a more general behavior, and the other forms are the rest of the six failure modes you will meet in 0.3.

It reaches for the common column when several could hold what you want, for the common table when the events live somewhere less popular, and for the common join key when the obvious field is not unique.

IT REACHED FOR        INSTEAD OF                    AND THE RESULT
 
TimeGenerated         the table's own event time    rows, wrong window
SigninLogs            AADNonInteractive...          rows, wrong table
UserPrincipalName     a genuinely unique key        rows, inflated count
 
ALL THREE RUN. ALL THREE RETURN ROWS.

Why there are six failure modes and not sixty. They are six expressions of a single mechanism: probable beats correct wherever the two come apart. That is why the list is short enough to memorize, why it did not change when the models got better, and why it will not change again. A more capable model is more often right about which value you meant. It is producing probable text either way.

Testing the rule on a table you do not know

The point of a mechanism rather than a list is that it works on material nobody prepared for you. Try it on a table this course has not introduced.

DeviceProcessEvents records process creation on endpoints. Suppose you ask for processes that were not started by a signed binary. Before reading further, work out which clause is at risk.

YOUR REQUEST   "processes not started by a signed binary, last 24 hours"
 
WHAT CAME BACK
  DeviceProcessEvents
  | where Timestamp > ago(24h)
  | where InitiatingProcessSignerType == "Microsoft"
  | project DeviceName, FileName, InitiatingProcessFileName
 
WHICH CLAUSE DID YOUR WORDING LOSE?

The clause at risk is the signature condition: almost every process on a working Windows host is signed, so almost every documented example is. You have just predicted a failure in a table you have never queried.

What it does not explain

Two things sit outside it. For both, the check is not a check on the query.

Estate facts. Nothing tells the model that 10.0.1.x is Manchester or that svc-legacyerp authenticates nightly by design. This is not the wrong token winning: the answer was never available, and the check is to ask your own data.

Counting. When a summary reports "fourteen files accessed", no counting occurred. A plausible number is more probable than an implausible one, which is why the wrong figures are the believable ones.

WHAT THE CHECK TOUCHES
 
inside the mechanism      the QUERY          read the clause
  plausible field                             against the schema
  silent window
  wrong join key
 
outside the mechanism     the WORLD          you cannot settle
  an estate fact                              this by reading
  a generated figure                          anything on screen

The two are not equally expensive. An estate fact costs one query and you own the answer permanently. A generated figure costs you the work of finding where it came from, and if no query produced it there is nothing to find, which is itself the answer. Both appear in 0.3 as named modes, and neither can be settled by reading. That is the practical division to carry out of this section: some of what comes back is wrong in the query, where you can see it, and some is wrong about the world, where you cannot. The first kind is what the rest of this module trains. The second kind is why the next section is about the estate.

07

The check you can run from now on

Which clause had a convention competing with your request

Before running any generated query, one question, and you now know why it works:

Which clause had a convention competing with my request?

Read the query looking for fields where documentation has a default: result codes, status strings, time columns, table names, join keys. Those are where your wording lost. Everything else, the account name, the projection, the sort order, almost certainly carried your intent, because there was no convention to override it.

On the scenario query that takes about ten seconds and lands on ResultType == 0 before you have run anything.

EmailEvents
| where RecipientEmailAddress == "r.scott@ne.com"
| where DeliveryAction == "Delivered"
| summarize Blocked = count() by SenderFromDomain
 
CONVENTION COMPETING?   DeliveryAction   YES, "Delivered" is the
                                         documented happy path
                        RecipientEmail   no, you supplied it
                        SenderFromDomain no, your projection
 
AND THE NAME DISAGREES WITH THE FILTER ABOVE IT.

That took about ten seconds and it did not require knowing the EmailEvents schema. You read the clauses, asked which of them had a documented default to compete with, and found one. Note what the check does not do: it does not tell you the query is wrong. DeliveryAction == "Delivered" may be exactly what you wanted. It tells you which clause to spend your attention on, and on this query the answer is one clause out of three.

Sometimes the answer is none. A query built entirely from values you supplied, with no result code, no status string, no time column and no join, has no contested position and the correct action is to run it.

SigninLogs
| where UserPrincipalName == "r.scott@ne.com"
| project TimeGenerated, IPAddress, AppDisplayName
| sort by TimeGenerated desc
 
RESULT CODES?   none      STATUS STRINGS?  none
TIME COLUMNS?   yours     JOINS?           none
 
NOTHING CONTESTED. RUN IT.

That case is common, and it is why the check is worth having rather than a general resolution to be careful. A check that fires on everything tells you nothing, and this one is silent on most of what you will see. The same reading works on a table you have never opened, which is the return on understanding the mechanism rather than memorizing a list.

08

Practice

Find the contested clause in a query of your own
hands on

The rule is worth nothing until you have seen what one contested clause costs on real data. Two runs of the same query, one number each.

Practice Find the contested clause
  1. Run the scenario query as written. Note the row count.
  2. Change ResultType == 0 to != 0. Run it again.
  3. Write down both numbers.
You now know what the gap between a documented default and your actual question costs, on one real account.

Then take a query somebody generated for you at work this week and read it once, looking only at the fields with a documented default. You are not checking it. You are finding out how many contested positions a real query has, which is usually two or three clauses out of ten.

Next: section 0.2 turns to why that error survives a careful reading, which is a fact about human attention rather than about the model, and is the reason a check has to be specific to be a check at all.