In this section

Catching Generated Query Errors: Your First Drill

Module 0

Introduction

Argument is the part of training that transfers least, and everything in this module up to here has been argument. This is where you find out whether it landed.

Four queries, generated from requests an analyst actually made, against the live estate. Each one runs and returns a plausible number of rows. Three are wrong in different ways and one is fine, and you are not told which until you have committed to a judgment. The drill measures two separate things: whether you catch the broken ones, and whether you can pass the sound one without flagging it. The second is the half nobody practices and the half that decides whether this habit survives a working month.

Scenario

You are covering a shift. Four queries have come back from an assistant in the last twenty minutes, each one produced from a request an analyst actually made. Every one runs. Every one returns rows. You have about a minute each and whatever you decide is what goes in the tickets.

Four generated queries. Each one runs. Each one returns rows. Three of them are wrong and one is fine, and the point of the fourth is that a habit which flags everything is as useless as one that flags nothing.

Work them in order. For each: read the request, read the query, decide before you run it, then run it and see whether the data agrees with you.

01

Before you start

Two rules that make the drill work

Two rules decide whether the drill teaches you anything, and both are about what you do before you run a query rather than after.

The four queries below were generated from the requests shown above them. They are realistic: each is the kind of query you get back in four seconds, each is syntactically valid, each references real columns in the Northgate corpus, and each returns rows. None of them fails in a way you can see from the output.

FOUR QUERIES. ALL FOUR RUN. ALL FOUR RETURN ROWS. DRILL 1 wrong DRILL 2 wrong DRILL 3 wrong DRILL 4 sound You are not told which is which until you have committed. The fourth is there because a habit that flags everything gets abandoned inside a fortnight.

You are not told which is which until you have committed to a verdict, which is the whole design of the exercise.

Two rules make this drill work.

Decide before you run. Write down, even as one word, whether you think the query is sound and which of the six modes it is exposed to. An unwritten judgment adjusts itself to the answer without your noticing, and you will finish believing you caught things you did not.

Run them. The blocks are editable and execute against the live corpus. Reading that a query returns 31 rows is a different experience from watching 31 rows appear when you expected none.

An unwritten judgment adjusts itself to the answer without your noticing, which is why the verdict goes down before the query runs.

The corpus is anchored in March 2026. Northgate's telemetry covers a thirty-day window ending 15 March 2026. That is a fact about this environment and it matters for the first drill, in exactly the way a fact about your own estate would matter for a query somebody generated for you. Estate context is not an abstraction; it is the difference between a right answer and a plausible one.

Reference How to run the drill
FOR EACH OF THE FOUR
 
  1  Read the REQUEST before the query. Most of the failures
     are visible in the gap between the two.
 
  2  Commit to a verdict before scrolling: sound, or which
     of the six modes is it exposed to?
 
  3  Run it. The result is often what makes the failure
     visible, and often not.
 
  4  Only then read the explanation.
 
SCORE TWO THINGS SEPARATELY
  how many of the three broken ones you caught
  whether you passed the sound one without flagging it
Committing before you scroll is the whole design. A verdict you formed after reading the answer measures nothing.

Work them in order and resist scrolling. The drill is short and the temptation to read ahead is the same temptation that closes an alert on a plausible answer, which is the habit it exists to test.

02

Drill 1

The outcome of every sign-in attempt over the last week
query to run

The first one is the request an analyst would actually make after an alert fires, phrased the way people phrase it under time pressure. Read the request before the query, because the gap between them is where the drill lives.

THE REQUEST, AS TYPED
 
  An alert fired on r.scott. Show me the outcome of every sign-in
  attempt on that account over the last week so I can see whether
  the attack succeeded.
SigninLogs
| where UserPrincipalName == "r.scott@ne.com"
| where TimeGenerated > ago(7d)
| summarize Attempts = count() by ResultType

Decide first. Sound, or which failure mode?

Run it. You get one row: ResultType 0, with 31 attempts. Thirty-one successful sign-ins, no failures at all, which reads as a quiet week on a healthy account.

The Reasonable Mistake

Reading a clean result as evidence the attack failed

This is a silent window. The corpus is anchored in March and ago(7d) is measured from now, so the window sits entirely after the incident. The eighty-four attempts on the night of 2 March, of which eighty-three failed, are outside it.

Nothing in a result of 31 suggests a window problem. Remove the time filter entirely and the same query returns 130 sign-ins with three distinct result codes. The filter was doing the work, not the analysis.

The check: move the boundary deliberately. If widening the window changes the answer's shape rather than its size, the boundary was the finding.

03

Drill 2

Which accounts that address caused failures on
query to run

The second changes one thing about the first: it names an address rather than an account. Watch what that does to the field the query has to filter on.

THE REQUEST, AS TYPED
 
  That Lithuanian address hit r.scott hard. Did it touch anyone
  else? Show me the failures it caused, by account.
SigninLogs
| where IPAddress == "45.83.64.117"
| where ResultType != 0
| summarize Failures = count() by UserPrincipalName

Decide first.

Run it. One row: r.scott@ne.com, 83 failures.

This one is correct, and that matters. The analyst asked which accounts the address caused failures on. The query filters to that address, excludes successes because failures were what was asked for, and groups by account. Every clause maps to something in the request.

The single row is a real finding: this address went after one account rather than spraying. If you flagged it as suspicious because the previous drill was wrong, you have learned something about your own calibration rather than about the query.

ONE ACCOUNT, TWO TABLES SigninLogs 4 failed interactive sign-ins reads as mild, unsuccessful pressure AADNonInteractiveUserSignInLogs 779 token-backed authentications where a replayed session actually lives A real number in the wrong table is the most dangerous shape in this course.

Same account, same night, two tables. The smaller number is the one that reads as reassuring.

04

Drill 3

Whether a stolen token was used to authenticate
query to run

The third is about a token rather than a password, which matters because a stolen session does not authenticate the way a stolen credential does. Ask yourself where that activity would be recorded before you look at the table it chose.

THE REQUEST, AS TYPED
 
  We think c.richardson's session token was stolen. Has that account
  authenticated at all since the phishing email?
SigninLogs
| where UserPrincipalName == "c.richardson@ne.com"
| where ResultType != 0
| summarize Attempts = count()

Decide first.

Run it. 4.

The Reasonable Mistake

Answering a question adjacent to the one asked

Two failures at once. The analyst asked whether the account authenticated; the query filters ResultType != 0, which counts only the attempts that FAILED. That is right answer, wrong question.

Worse, a stolen session token does not produce an interactive sign-in at all. Token replay is recorded in AADNonInteractiveUserSignInLogs, where this account has 779 events. Querying SigninLogs for token abuse is a confident absence waiting to happen: the wrong table returns a small number rather than an error.

The check: name the semantic claim. For this to answer whether the account authenticated, the filter would have to include successes, and the table would have to be the one that records this kind of authentication. Both fail.

05

Drill 4

A count of distinct users the policy did not apply to
query to run

The fourth returns a single number, which is the shape most likely to be pasted into a ticket unexamined. Decide what that number would have to count for the answer to be right.

THE REQUEST, AS TYPED
 
  Give me the count of distinct users that Conditional Access did
  not apply to, so I can size the gap.
SigninLogs
| where ConditionalAccessStatus == "notApplied"
| summarize Events = count()

Decide first.

Run it. 12.

The Reasonable Mistake

A count of the wrong thing, reported as the right thing

The analyst asked for distinct users. The query counts events. Twelve is a real number and it is not the number requested, and if it goes into a report as "twelve users affected" it is invented precision introduced by the analyst rather than by the model. Change count() to dcount(UserPrincipalName) and the answer changes.

The pattern across all four is easier to see laid out than described.

Test it on your own assistant

Try this Have it grade its own work
Write me a KQL query against SigninLogs that counts distinct users

who signed in from outside our corporate IP ranges.

Then list every assumption you made that could make this query wrong.

What to look at. The assumption list. It will be a good list, and it will usually name the real problem: that nothing in the request said what your corporate ranges are.

What this demonstrates. Asked directly, these systems are often able to name their own weak points, which is genuinely useful and easy to mistake for a solution. The list is generated the same way the query was, so it can be incomplete without any sign that it is.

Use it as a starting set of things to check rather than as a completed check. The four drills above are the same discipline: a plausible artifact, and one specific question that settles whether it holds.

06

Scoring yourself honestly

Detection and discrimination are separate skills

Before reading on, count. Not how many you flagged, but how you did on each of the two things separately, because they are different skills and most people are much better at one.

FILL THIS IN BEFORE YOU READ THE NEXT PARAGRAPH
 
  DETECTION      of drills 1, 3 and 4, how many did you mark
                 unsound BEFORE running them?          ___ / 3
 
  DISCRIMINATION did you pass drill 2 without flagging it?
                                                        yes / no
 
  NAMING         for each one you caught, did you name the
                 mode, or only feel that something was off?
                                                        ___ / 3
 
THREE NUMBERS. THE MIDDLE ONE IS THE ONE NOBODY DRILLS.

Detection. Of the three broken queries, how many did you mark as unsound before running them? This is the number people expect to be the important one and it is the easier half. Anyone primed by a module on failure modes is suspicious of everything they are shown in an exercise.

Discrimination. Did you pass drill 2? An analyst scoring three out of three on detection and failing this one has a habit that flags everything, and that habit gets abandoned.

Catching all three and flagging the fourth is a worse result than catching two and passing the fourth.

Naming. For each one you caught, did you name the mode or just feel that something was off? "The window looks wrong" and "this is a silent window, so I widen the boundary and see if the shape changes" are different states. The first is unease and the second is a check with an end.

If this is the first time you have measured any of that on yourself, the number is more useful than any individual answer. It tells you which half of the skill your next month should go on.

THE NUMBER WORTH KNOWING ABOUT YOURSELF
 
  artifacts checked this week          ____
  problems found                       ____
  your hit rate                        ____%
 
Under 10%   your attention belongs on the pre-flight in 0.5,
            not on checking. Prevent rather than catch.
10 to 30%   normal. The sweep is earning its time.
Over 30%    something about how you are asking is generating
            failures. Look at the requests, not the answers.

Whichever band you land in, the number is about your requests and your tables rather than about the tool, so it moves when you change what you do. That is the only reason it is worth writing down.

07

What each one cost you

What a miss would have done on a real shift

Worth walking back through them, because the interesting number is not how many you caught but what each miss would have done on a real shift.

Reference What each miss costs on a real shift
DRILL  THE MISS                       WHAT IT COSTS
──────────────────────────────────────────────────────────────
1      window excludes the event      the alert closes as
                                      "no activity found"
 
2      filters to the wrong entity    the address is cleared
                                      while the account is not
 
3      wrong table for the behavior   token replay invisible;
                                      the session survives
                                      the password reset
 
4      a number that counts something you report a figure you
       adjacent to the question       cannot defend when asked
──────────────────────────────────────────────────────────────
Three of the four end with somebody writing "no evidence found".
None of the four is a wrong answer. All four are correct queries answering a question next to the one asked.

Drill 1 closes the alert. An analyst who accepts thirty-one clean successes writes "no failed authentication observed in the period" and moves on. The brute force is a day outside the window, the account was compromised at 22:55 on 2 March, and nothing in the ticket records that the period examined was chosen by a default rather than by the incident. Six months later nobody can tell that the wrong week was searched, because the query is in the ticket and it looks correct.

Drill 2 costs you nothing, and flagging it costs you something. An analyst who marks it suspicious has not been careful, they have been indiscriminate, and a colleague who queries every generated artifact is slower than one who writes their own queries, which is the outcome that gets the tool abandoned.

Passing a sound artifact confidently is half the skill, and it is the half nobody drills because it feels like doing nothing.
DRILL 1   asked: did the attack succeed?
          answered: what happened in the last 7 days
          the compromise is 11 days old
 
DRILL 2   asked: which accounts did that address hit?
          answered: what did r.scott do?
          the address touched five accounts

Drill 3 produces a false negative on an active compromise. Four failed attempts on an account whose session token was stolen reads as an account under mild, unsuccessful pressure. The token replay is in a different table entirely, and the number four is small enough to be reassuring rather than small enough to be suspicious. That combination, a real number in the wrong table, is the most dangerous shape in this course.

Drill 4 puts a wrong figure in a report. Twelve is a real count of real events, and reported as twelve affected users it overstates by one. Nobody dies of that. But the analyst has now published a number they did not measure, and the next figure they publish will be trusted on the same basis, which is the mechanism by which invented precision leaks from a summary into an organization's understanding of itself.

WHAT A MISS ACTUALLY COSTS Drill 1 silent window Alert closed. Account was compromised at 22:55, one day outside the window. Drill 2 sound Flagging it costs you speed, and eventually costs the team the tool. Drill 3 wrong question and wrong table False negative on an active compromise. Four is small enough to reassure. Drill 4 invented precision A number you did not measure enters a report, and the next one is trusted too.

Three different costs, and only one of them is a missed detection. Flagging a correct artifact costs the team the tool.

08

The minute you actually have

One check each rather than five, and how it is chosen

One practical note before the scoring, because the drill gives you unlimited time and a shift does not.

On a real queue you will not run five checks on four artifacts. You will run one check on each, chosen from what the request contained, and you will be right often enough that the habit survives. Drill 1 needed only the time boundary widened. Drill 3 needed only the question "does this table hold the kind of authentication I am asking about". Drill 4 needed only "is this counting the thing I asked for".

ONE CHECK EACH, CHOSEN FROM THE REQUEST
 
DRILL 1   request said "around the alert"
          -> widen the window a day each way          ~20 sec
 
DRILL 2   request was absolute and single-table
          -> nothing. Pass it.                          0 sec
 
DRILL 3   request asked about token-backed sign-in
          -> does this table hold that?                ~30 sec
 
DRILL 4   request asked for affected USERS
          -> is this counting users or events?         ~20 sec
 
THREE OF FOUR, UNDER TWO MINUTES, AND NO SWEEP WAS RUN.
One check each, chosen deliberately, catches three of the four in under two minutes total.

That is the working target, and it is why the sweep in 0.3 is ordered by cost rather than by mode number.

09

Why three wrong and one right

Why an all-broken exercise set teaches the wrong reflex

The ratio is deliberate and it is not a trick.

An exercise set where everything is broken teaches a reflex that is worse than useless: answer "wrong" without looking, be right most of the time, and learn nothing that survives contact with a tool that is correct most of the time. Real generated output is mostly fine. Six failure modes across dozens of queries a week is a minority of artifacts, and a habit calibrated for a majority will be discarded within a fortnight because it costs more than it returns.

WHAT AN ALL-BROKEN EXERCISE SET TEACHES
 
  strategy      answer "wrong" without looking
  score         100%
  transfers to  nothing. Real output is mostly fine, so the
                same strategy scores badly the first week
                you use it on a real queue
 
WHAT THIS SET TEACHES
 
  strategy      read the request, name the check, decide
  score         lower, and it is the score that survives

What you are building is not suspicion. It is discrimination, which is a harder thing: the ability to look at four artifacts and spend your minute on the one that needed it.

TWO SKILLS, AND THE SECOND IS THE ONE THAT LASTS DETECTION Did you catch the three broken ones? The easier half. Anyone primed by a module on failure modes is suspicious. DISCRIMINATION Did you PASS the sound one? The harder half. Nobody drills it, because it feels like doing nothing. 3 of 3 detected, sound one flagged → a habit that dies within a fortnight 2 of 3 detected, sound one passed → a habit you still have in a year

Catching the broken ones is the half a module on failure modes primes you for. Passing the sound one is the half that decides whether you still do this in a year.

That is why the drill includes a correct query, why the module summary asks you to say why it was correct, and why every module after this one includes sound artifacts alongside broken ones without telling you the proportion.

The proportion in this drill is not the proportion in your estate, and it would be a mistake to learn a number from it. Three broken in four is a teaching ratio, chosen so the sound query has somewhere to hide.

What transfers is the procedure rather than the frequency: read the request, name the check that fits it, run the check, decide.

If you catch yourself estimating how many are usually wrong before you look, you have gone back to guessing, and guessing is the thing the check replaces.

10

What the drill was measuring

Not suspicion. Discrimination.

Not whether you spotted three errors. Anyone primed by a module on failure modes will be suspicious of everything.

What it measured is whether you passed drill 2. An analyst who flags the correct query has a habit that will be abandoned within a fortnight, because a colleague who distrusts every generated answer has stopped getting value from the tool and everyone around them knows it.

The working target is not maximum suspicion. It is naming which of the six a query is exposed to, running the one check that settles it, and moving on in under a minute.

Reference The sweep, on any generated query ~60 seconds
1  Does the variable name agree with the filter above it?
2  Is the window absolute, and is the anchor an EVENT time?
3  Is this the table where that behavior is recorded here?
4  If it returned nothing, can it return rows at all?
5  Is any number in it counting what the question asked?
 
Then: name which of the six it is exposed to, run the one
check that settles it, and move on. Including when the
answer is "this one is fine".
Five checks, and the fifth is the one that catches a figure on its way into a report.

That card is the module in one page. Everything after this is the same five questions asked of harder artifacts: a detection rule instead of a query, an incident summary instead of a row count, a runbook instead of a filter.

11

Practice

Defend the correct one, then bring an artifact of your own
hands on

Four drills is a measurement of one afternoon. The habit is what you do on Monday, and it takes about ten minutes to start.

Practice Defend the correct one
  1. Go back to drill 2, the query that was sound.
  2. Write one sentence saying why it is sound. Not "it looked fine": name the reason.
  3. Check your sentence against the three clauses: the filter, the grouping, the request.
Being able to say why a correct artifact is correct is the half of this skill nobody practices, and it is what stops you flagging everything.

Then bring one artifact of your own: a query, a rule, or a summary an assistant produced for you at work. Run the sweep on it and name which of the six it was exposed to. The course provides the drill; your own estate provides the case where it matters.

Next: the module summary collects what this module established and gives you a first step for your next shift.