In this section

Why an AI Assistant Cannot Know Your Environment

Module 0

Introduction

Every exercise in this course runs against one fictional estate, and this section introduces it. Not as a tour of a company that does not exist, but because the checks in every later module resolve to a question about data, and you cannot judge whether an answer is plausible without knowing what ordinary looks like in the place it came from.

That knowledge is also the thing an assistant can never have. This section draws the line precisely: which kinds of context you can hand over in ten lines at the start of a session, and which kinds exist only as expectations you built by working somewhere. By the end you will know the tables you will be querying, the four intrusions hidden in the traffic, and why the ordinary traffic matters more than the intrusions do.

Scenario

A generated query returns three rows for a service account you have not seen before. The assistant summarizes them accurately. To decide whether three rows is ordinary or alarming you need to know what that account normally does, and that fact exists in exactly one place: the estate's own history. This section introduces the estate you will use for every exercise, and explains why the knowledge it gives you cannot be supplied any other way.

01

Why a fixed estate

One environment across every exercise, and what that buys

Most security training that involves AI stops at description. You read about a failure mode, you nod, and you never encounter one under conditions where you might miss it. That produces recognition without competence: you can define the six failure modes and you still approve the query.

The Reasonable Mistake

Taking recognition for competence

Reading a worked example of a failure and finding it obvious is the normal experience, and it predicts almost nothing about catching the same failure at alert nineteen with the answer already written in the ticket. The example was labeled. The one in front of you is not, and the labeling is most of what made it visible.

Verification cannot be taught that way. The skill is not knowing that generated output can be wrong, since everyone knows that already.

The skill is catching it in a specific case where it is not obvious, and that needs a case with data behind it that you can interrogate.

So every exercise in this course runs against a single fixed environment: Northgate Engineering, an 810-person firm whose telemetry sits behind the query surface on these pages. When a module shows you a generated query, you run it. Real rows come back. When the finding tells you the query was wrong, you can prove that for yourself by changing one clause and running it again.

SigninLogs
| summarize SignIns = count() by bin(TimeGenerated, 1d)
| sort by TimeGenerated asc

Run it before reading on. Nothing about that query is interesting, which is the point: the daily volume it returns is the first thing you now know about this estate that you did not know a minute ago, and every later judgment about what is unusual rests on a number like it.

Several failure modes are only visible to somebody who knows what normal looks like, which is why the estate does not change between modules.

A service account on a laptop is suspicious because you know service accounts touch servers. Three office addresses are reassuring only if you know the office range.

WHAT YOU WILL KNOW ABOUT THIS ESTATE BY MODULE 5
 
  where sign-ins normally originate
  which hosts service accounts normally reach
  what a normal daily failure count looks like
  which hosts run PowerShell as a matter of routine
 
ALL FOUR ARE EXPECTATIONS RATHER THAN FACTS, AND BEING SURPRISED
BY ONE IS THE CONDITION UNDER WHICH VERIFICATION WORKS AT ALL.

None of those four is written down anywhere. You acquire them the way you did at work, by looking at the same place repeatedly, which is why the estate does not change between modules.

02

The tables you will use

Identity, endpoint, email and cloud, and what each records

The corpus holds 21 tables and roughly 79,000 rows, of which this course uses a handful heavily.

Identity, which is where most of the verification exercises live because sign-in data is where the subtle field errors are:

  • SigninLogs: interactive sign-ins, with ResultType, IPAddress, AppDisplayName, ConditionalAccessStatus and AuthenticationRequirement
  • AADNonInteractiveUserSignInLogs: the token refreshes and background authentications that a query filtered to SigninLogs will silently miss
  • AuditLogs: directory changes: role assignments, method registrations, application consent
  • IdentityLogonEvents: on-premises domain authentication, Kerberos and NTLM

Endpoint, from Defender:

  • DeviceProcessEvents: process creation with command lines and parent process
  • DeviceNetworkEvents, DeviceFileEvents, DeviceRegistryEvents, DeviceLogonEvents, DeviceEvents

Email and cloud: EmailEvents, OfficeActivity, CloudAppEvents.

Alerts: SecurityAlert, AlertEvidence, SecurityIncident.

The Reasonable Mistake

Two pairs of tables that produce failure mode 4

SigninLogs and AADNonInteractiveUserSignInLogs both hold sign-ins, and a token replay lives in the second. DeviceLogonEvents and IdentityLogonEvents both hold logons, and domain authentication lives in the second. In both cases, querying the wrong one of the pair returns an empty result rather than an error, and an empty result reads as an answer. You will meet both in Module 3.

Four of these tables come in two pairs that look interchangeable and are not, and knowing which of a pair holds what you are asking about is a fact about this estate that no assistant can supply.

03

Four intrusions, and why they matter here

What is hidden in the traffic, and why it is subtle

Four attack chains run through the corpus. A generated query that fails silently is only dangerous if there was something to find, and these are the something.

Adversary-in-the-middle phishing. A proxy captures a session token from c.richardson@ne.com and the attacker uses it from outside the estate. The sign-in satisfies MFA because the token already carries the claim, so a query looking for failed authentication finds nothing.

Password spray. One address makes 21 attempts across 20 accounts, one each, so no account locks and no per-account threshold fires. A query grouped by user shows one failure per user and reads as ordinary noise.

WHY THE ORDINARY TRAFFIC IS THE POINT ~79,000 ROWS OF ORDINARY ACTIVITY people signing in, service accounts on schedule, PowerShell where it always runs four intrusions, buried in it A wrong query returns rows because there are always ordinary rows to return. That condition cannot exist in a dataset containing only attacks.

A wrong query returns rows because ordinary rows outnumber everything else.

Endpoint compromise and credential theft. Two laptops are dumped with comsvcs.dll MiniDump, a signed Microsoft binary. Antivirus has nothing to say about it and neither does a query looking for unusual executables.

Ransomware pre-encryption. Shadow copy deletion and backup interference before any encryption, which is the window in which it is still preventable.

Every one of the four is built so that the obvious query returns nothing, or something that reads as ordinary.

Alongside these sits a brute force against r.scott@ne.com: 84 attempts from one address on the night of 2 March, 83 of them failed and one succeeded at 22:55:12. That single success is the row that most of the query exercises in this course are built around, because it is easy to write a query that runs, returns rows, and does not contain it.

THE ROW MOST OF THIS COURSE IS BUILT AROUND
 
TimeGenerated          2026-03-02T22:55:12Z
UserPrincipalName      r.scott@ne.com
IPAddress              45.83.64.117
ResultType             0
 
IT IS THE ONE SUCCESS AMONG 84 ATTEMPTS, AND ResultType 0 IS WHY
A QUERY WRITTEN TO FIND "FAILURES" RETURNS IT AND NOTHING ELSE.

Hold that row in mind. Every drill in 0.7 is a query that runs, returns a reasonable number of rows, and either contains it or does not.

04

The noise is the point

Why a wrong query still returns rows

An estate containing only attacks would teach nothing, because every query would find the intrusion and every verification would pass. This corpus is mostly ordinary: people signing in from the office, service accounts on schedule, developers running PowerShell for legitimate reasons, someone in finance downloading eleven files because that is their job.

That ordinary activity is what makes verification a skill rather than a formality:

  • A query that accidentally selects successes still returns plenty of rows, because there are plenty of successes
  • A time window one day off still finds failures, because there are always a few failures
  • A rule matching PowerShell execution fires constantly, because PowerShell runs constantly
THE WRONG QUERY, AGAINST A CORPUS OF ONLY ATTACKS
  returns nothing, and you notice immediately
 
THE WRONG QUERY, AGAINST THIS CORPUS
  returns rows, because there are always ordinary rows
  to return, and the rows look like an answer
 
THAT SECOND CONDITION IS THE ONE YOU WORK IN.

In every case the wrong query produces a confident, populated, plausible answer, and that condition cannot be reproduced in a dataset containing only the attack.

What "normal" looks like here, and why you need it

Verification requires a baseline, and a baseline is not a fact you can look up. It is a set of expectations you build by looking at the estate when nothing is wrong. Four are worth acquiring now, because exercises later depend on them.

Keep this Northgate's baseline, in four numbers
WHERE SIGN-INS COME FROM
  10.0.1.x is Manchester. Bristol has its own range.
  Anything else is a remote worker, a partner VPN egress,
  or an intrusion, and which one is a question about the
  ACCOUNT rather than about the address.
 
WHEN SERVICE ACCOUNTS RUN
  33 of 36 svc- logons go to SRV-NGE-MCR hosts, on schedule.
  The three that do not are an exercise later in this course.
 
WHAT A NORMAL FAILURE RATE LOOKS LIKE
  A handful a day. People mistype passwords. "Some failures"
  tells you nothing; the brute force is identified by SHAPE,
  one address against one account 84 times.
 
HOW MUCH NOISE POWERSHELL MAKES
  17 hosts, constantly, legitimately. Any rule that starts
  "alert on PowerShell" dies on contact with this estate.
Four questions, not four facts. Yours have different answers and you get them the same way, by running the query when nothing is wrong.

The baseline is what a model cannot have. Every fact in that card is a property of this estate. None of it is in any model's training data and no amount of prompting will surface it.

The assistant supplies fluency with the language. You supply everything about the place.

Where a generated answer goes wrong, it is very often because it applied a general truth to a specific estate where the general truth does not hold.

A note on how the corpus was built

The telemetry is generated rather than captured, and the way it was generated matters for what you can conclude from it.

Ordinary activity was simulated: users with roles, working patterns, devices and habits, producing the volume and rhythm a real estate produces. The attacks were then run through that activity as chains, so the malicious events sit inside the noise in the right proportion rather than being appended to it.

HOW THE CORPUS WAS ASSEMBLED
 
  1  users given roles, hours, devices and habits
  2  ordinary activity simulated from those, at real volume
  3  the four attack chains RUN THROUGH that activity
 
NOT: attacks generated separately and appended. The order in
step 3 is why a wrong query returns ordinary rows rather than
returning nothing.

The practical consequence is that absence is meaningful here. If a query finds nothing, the reason is either that the thing did not happen or that you looked in the wrong place, and both are real answers. The corpus is not a puzzle where every query is supposed to return something interesting, and treating it as one will cost you a finding.

05

The same log line, twice

An accurate translation, and a finding

Here is a single event, presented without comment.

THE SAME LINE, READ TWICE svc-sql NE-LEWIS-LT Network Ntlm Success 2026-03-12T22:08:47Z WITHOUT THE ESTATE A service account authenticated to a host over NTLM. It succeeded. Accurate. A translation of the record. WITH THE ESTATE A database account on a LAPTOP, over the old protocol, once in thirty days. A finding. The difference is not intelligence or prompting. The second reader has seen the other thirty-five logons.

Both readings are accurate. Only the second is a finding, and the difference is the thirty-five other logons that reader has seen.

2026-03-12T22:08:47Z  svc-sql  NE-LEWIS-LT  Network  Ntlm  Success

Ask an assistant what it means and you will get an accurate reading: a service account named svc-sql authenticated to a host called NE-LEWIS-LT over the network using NTLM, and it succeeded. Every word of that is correct.

Now the same line read by someone who knows Northgate. svc-sql is a database service account. NE-LEWIS-LT is a laptop. Service accounts in this estate authenticate to SRV-NGE-MCR servers, thirty-three times out of thirty-six. NTLM rather than Kerberos is the older protocol and this estate uses it for legacy applications on servers, not for laptops. And svc-sql appears exactly once in thirty days of logon history, which is this line.

The first reading is a translation of the record. The second is a finding, and the gap between them is thirty-five logons the second reader has seen.

What "context" turns out to mean. People say a model lacks context as though context were a paragraph that could be supplied. In this example the load-bearing context is a distribution: what these accounts usually do, expressed as a ratio nobody has ever written down. It exists in the analyst's head as an expectation, and it got there by looking at the estate over months.

Test it on your own assistant

Try this Ask it what normal looks like
In a Microsoft 365 environment with about 800 staff, how many failed

sign-ins would you expect on an ordinary weekday, and how many would indicate a password spray attack?

What to look at. How specific the answer is. You will very likely get numbers, or a range, delivered with the same confidence as a fact. Nothing about your environment was in the question.

What this demonstrates. The answer is a plausible generalization, and watching one arrive is the useful part.

WHAT IT TOLD YOU        a typical estate sees a handful of failed
                        sign-ins per user per day
 
WHERE THAT CAME FROM    the question asked for a number, so a
                        number was produced
 
WHAT SETTLES IT         one summarize over your own SigninLogs
 
YOUR ANSWER DEPENDS ON YOUR MFA POSTURE, YOUR LOCKOUT POLICY AND
HOW MANY SERVICE ACCOUNTS YOU RUN. NONE OF THAT WAS AVAILABLE.

Numbers arrive because numbers are what the question asked for, not because anything was measured, and your own estate settles it in one query.

06

Four kinds of context

Which can be supplied to an assistant and which cannot

It helps to separate them, because they differ in whether they can be supplied.

Stated facts. The estate has 810 staff, two offices, an M365 E5 tenant, a Palo Alto perimeter. Writeable, short, easily supplied.

Conventions. Hosts are named SITE-FUNCTION-NN. Service accounts are prefixed svc-. Office ranges are 10.0.1.x and Bristol's equivalent. Writeable with a little effort, and genuinely useful when supplied.

Distributions. How often service accounts touch laptops. What a normal daily failure count looks like. How many files finance downloads at month end. Not writeable, because they are not facts you hold as sentences. You hold them as expectations, and you discover you hold them when something violates one.

CAN YOU WRITE IT DOWN?
 
  Stated facts        yes, in one line     810 staff, E5, two sites
  Conventions         yes, with effort     SRV- is a server
  Distributions       NO                   you hold these as
                                           expectations, not sentences
  Exceptions          partly, never fully  the list is generated by
                                           events rather than design
 
THE TWO YOU CANNOT WRITE DOWN ARE THE TWO THE FINDINGS COME FROM.

Exceptions and history. That svc-legacyerp authenticates without MFA because of an application nobody has rewritten, that a partner's VPN egress looks foreign every Tuesday, that the alert on this host last month turned out to be a misconfigured backup agent. Partly writeable and never complete, because the list is generated by events rather than by design.

FOUR KINDS OF CONTEXT. ONLY TWO CAN BE SUPPLIED. WRITEABLE, AND WORTH WRITING Stated facts 810 staff, M365 E5, Sentinel, Defender XDR Conventions WHERE THE FINDINGS ARE Distributions 33 of 36 service logons go to servers Exception history Ten lines pasted at the start of a session measurably improves query generation You do not know what you know until something violates it The assistant supplies fluency with the language. You supply everything about the place. That division does not move.

Two of the four fit in ten lines at the start of a session. The other two are why an assistant cannot reach a finding on its own.

The first two can be given to an assistant. The second two are where the findings are, and they resist being written down.

07

What you can actually supply

Ten lines of preamble, and what it will not fix

Supplying the writeable half produces a real improvement, and three things carry most of it.

The naming conventions, because they let an assistant read a hostname as structured rather than as a string. Told that SRV- means server and NE-SURNAME-LT means a user laptop, the shape of the anomaly becomes visible from the names alone.

ONE LINE OF PREAMBLE, AND WHAT IT CHANGES
 
YOU WRITE   "SRV- is a server. NE-SURNAME-LT is a user laptop."
 
BEFORE      "svc-sql authenticated to NE-LEWIS-LT"
            a hostname, treated as a string
 
AFTER       "a service account authenticated to a USER LAPTOP"
            a shape, and the shape is the anomaly

The address ranges, because internal and external is the most common distinction in security reasoning and it is entirely estate-specific.

The schema is the single highest-value thing to supply, because it attacks failure modes 1 and 4 directly. A model told which of the two logon tables holds domain authentication is much less likely to query the wrong one.

SAME REQUEST: "did this service account touch any laptops?"
 
WITHOUT THE PREAMBLE   queried DeviceLogonEvents, guessed at what
                       a laptop is, matched on a name pattern it
                       invented
 
WITH IT                queried IdentityLogonEvents, used NE-*-LT
                       because you said that is a user laptop
 
WHAT CHANGED           the table and the convention. NOT whether
                       three logons to a laptop is normal here.

A reusable preamble is worth writing once. Ten lines describing conventions, ranges and the tables you use, pasted at the start of a session, measurably improves query generation and costs nothing after the first time. It does not make the assistant know your estate; it makes it stop guessing about the parts you could have told it.

The last line is the boundary. Ten lines of preamble fixed which table to query and how to read a hostname, and it did not touch the only question that decides whether three logons matter. That question is yours, and it stays yours no matter how much context you supply.

08

Running your first query

A baseline against the estate
query to run

Confirm the mechanism before an exercise depends on it. The block below is editable and runs against the live corpus.

SigninLogs
| where IPAddress == "45.83.64.117"
| summarize Attempts = count() by ResultType
| sort by Attempts desc

You should get three rows: 50126 with 58, 50053 with 25, and 0 with 1. Wrong password, account locked, and one success.

One row out of 84 is the entire finding, and it sits at the bottom of a sorted result.

Every exercise in Module 3 is a variation on a query that returns something reasonable while excluding that row.

If the block did not run. The query surface needs a signed-in account with an active subscription. If you are reading the free preview, the exercise blocks in later modules will not execute, and the reasoning in each one still stands: every generated artifact in this course is accompanied by the finding it produces, so you can follow the argument without running anything. You will get considerably more out of it running the queries.

09

Practice

Build a baseline for your own environment
hands on

You cannot borrow Northgate's baseline. The four expectations above are worth something only because somebody ran the queries that produced them, and the same four questions have different answers in your estate.

Practice Build a baseline in twenty minutes
  1. How many sign-ins on an ordinary day?
  2. From how many distinct addresses?
  3. Which accounts authenticate most?
  4. What does the daily failure count look like?
  5. Which hosts run PowerShell, and how often?
None of these is an investigation. All of them are what makes an investigation possible.

Answer them against Northgate, then against your own estate, and write the answers somewhere you will find them. When a generated query returns a number, the only thing that tells you whether it is plausible is knowing what normal looks like, and normal is not something you can look up.

Next: section 0.5 turns the six modes into a prediction you can make from your own request, before you have read a single row of output.