Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
Why an AI Assistant Cannot Know Your Environment
Introduction
Every exercise in this course runs against one fictional estate, and this section introduces it. Not as a tour of a company that does not exist, but because the checks in every later module resolve to a question about data, and you cannot judge whether an answer is plausible without knowing what ordinary looks like in the place it came from.
That knowledge is also the thing an assistant can never have. This section draws the line precisely: which kinds of context you can hand over in ten lines at the start of a session, and which kinds exist only as expectations you built by working somewhere. By the end you will know the tables you will be querying, the four intrusions hidden in the traffic, and why the ordinary traffic matters more than the intrusions do.
Scenario
A generated query returns three rows for a service account you have not seen before. The assistant summarizes them accurately. To decide whether three rows is ordinary or alarming you need to know what that account normally does, and that fact exists in exactly one place: the estate's own history. This section introduces the estate you will use for every exercise, and explains why the knowledge it gives you cannot be supplied any other way.
Why a fixed estate
One environment across every exercise, and what that buysMost security training that involves AI stops at description. You read about a failure mode, you nod, and you never encounter one under conditions where you might miss it. That produces recognition without competence: you can define the six failure modes and you still approve the query.
The Reasonable Mistake
Taking recognition for competence
Reading a worked example of a failure and finding it obvious is the normal experience, and it predicts almost nothing about catching the same failure at alert nineteen with the answer already written in the ticket. The example was labeled. The one in front of you is not, and the labeling is most of what made it visible.
Verification cannot be taught that way. The skill is not knowing that generated output can be wrong, since everyone knows that already.
The skill is catching it in a specific case where it is not obvious, and that needs a case with data behind it that you can interrogate.
So every exercise in this course runs against a single fixed environment: Northgate Engineering, an 810-person firm whose telemetry sits behind the query surface on these pages. When a module shows you a generated query, you run it. Real rows come back. When the finding tells you the query was wrong, you can prove that for yourself by changing one clause and running it again.
SigninLogs
| summarize SignIns = count() by bin(TimeGenerated, 1d)
| sort by TimeGenerated asc
Run it before reading on. Nothing about that query is interesting, which is the point: the daily volume it returns is the first thing you now know about this estate that you did not know a minute ago, and every later judgment about what is unusual rests on a number like it.
Several failure modes are only visible to somebody who knows what normal looks like, which is why the estate does not change between modules.
A service account on a laptop is suspicious because you know service accounts touch servers. Three office addresses are reassuring only if you know the office range.
WHAT YOU WILL KNOW ABOUT THIS ESTATE BY MODULE 5
where sign-ins normally originate
which hosts service accounts normally reach
what a normal daily failure count looks like
which hosts run PowerShell as a matter of routine
ALL FOUR ARE EXPECTATIONS RATHER THAN FACTS, AND BEING SURPRISED
BY ONE IS THE CONDITION UNDER WHICH VERIFICATION WORKS AT ALL.
None of those four is written down anywhere. You acquire them the way you did at work, by looking at the same place repeatedly, which is why the estate does not change between modules.
The tables you will use
Identity, endpoint, email and cloud, and what each recordsThe corpus holds 21 tables and roughly 79,000 rows, of which this course uses a handful heavily.
Identity, which is where most of the verification exercises live because sign-in data is where the subtle field errors are:
SigninLogs: interactive sign-ins, withResultType,IPAddress,AppDisplayName,ConditionalAccessStatusandAuthenticationRequirementAADNonInteractiveUserSignInLogs: the token refreshes and background authentications that a query filtered toSigninLogswill silently missAuditLogs: directory changes: role assignments, method registrations, application consentIdentityLogonEvents: on-premises domain authentication, Kerberos and NTLM
Endpoint, from Defender:
DeviceProcessEvents: process creation with command lines and parent processDeviceNetworkEvents,DeviceFileEvents,DeviceRegistryEvents,DeviceLogonEvents,DeviceEvents
Email and cloud: EmailEvents, OfficeActivity, CloudAppEvents.
Alerts: SecurityAlert, AlertEvidence, SecurityIncident.
The Reasonable Mistake
Two pairs of tables that produce failure mode 4
SigninLogs and AADNonInteractiveUserSignInLogs both hold sign-ins, and a token replay lives in the second. DeviceLogonEvents and IdentityLogonEvents both hold logons, and domain authentication lives in the second. In both cases, querying the wrong one of the pair returns an empty result rather than an error, and an empty result reads as an answer. You will meet both in Module 3.
Four of these tables come in two pairs that look interchangeable and are not, and knowing which of a pair holds what you are asking about is a fact about this estate that no assistant can supply.
Four intrusions, and why they matter here
What is hidden in the traffic, and why it is subtleFour attack chains run through the corpus. A generated query that fails silently is only dangerous if there was something to find, and these are the something.
Adversary-in-the-middle phishing. A proxy captures a session token from c.richardson@ne.com and the attacker uses it from outside the estate. The sign-in satisfies MFA because the token already carries the claim, so a query looking for failed authentication finds nothing.
Password spray. One address makes 21 attempts across 20 accounts, one each, so no account locks and no per-account threshold fires. A query grouped by user shows one failure per user and reads as ordinary noise.
A wrong query returns rows because ordinary rows outnumber everything else.
Endpoint compromise and credential theft. Two laptops are dumped with comsvcs.dll MiniDump, a signed Microsoft binary. Antivirus has nothing to say about it and neither does a query looking for unusual executables.
Ransomware pre-encryption. Shadow copy deletion and backup interference before any encryption, which is the window in which it is still preventable.
Every one of the four is built so that the obvious query returns nothing, or something that reads as ordinary.
Alongside these sits a brute force against r.scott@ne.com: 84 attempts from one address on the night of 2 March, 83 of them failed and one succeeded at 22:55:12. That single success is the row that most of the query exercises in this course are built around, because it is easy to write a query that runs, returns rows, and does not contain it.
THE ROW MOST OF THIS COURSE IS BUILT AROUND
TimeGenerated 2026-03-02T22:55:12Z
UserPrincipalName r.scott@ne.com
IPAddress 45.83.64.117
ResultType 0
IT IS THE ONE SUCCESS AMONG 84 ATTEMPTS, AND ResultType 0 IS WHY
A QUERY WRITTEN TO FIND "FAILURES" RETURNS IT AND NOTHING ELSE.
Hold that row in mind. Every drill in 0.7 is a query that runs, returns a reasonable number of rows, and either contains it or does not.
The noise is the point
Why a wrong query still returns rowsAn estate containing only attacks would teach nothing, because every query would find the intrusion and every verification would pass. This corpus is mostly ordinary: people signing in from the office, service accounts on schedule, developers running PowerShell for legitimate reasons, someone in finance downloading eleven files because that is their job.
That ordinary activity is what makes verification a skill rather than a formality:
- A query that accidentally selects successes still returns plenty of rows, because there are plenty of successes
- A time window one day off still finds failures, because there are always a few failures
- A rule matching PowerShell execution fires constantly, because PowerShell runs constantly
THE WRONG QUERY, AGAINST A CORPUS OF ONLY ATTACKS
returns nothing, and you notice immediately
THE WRONG QUERY, AGAINST THIS CORPUS
returns rows, because there are always ordinary rows
to return, and the rows look like an answer
THAT SECOND CONDITION IS THE ONE YOU WORK IN.
In every case the wrong query produces a confident, populated, plausible answer, and that condition cannot be reproduced in a dataset containing only the attack.
What "normal" looks like here, and why you need it
Verification requires a baseline, and a baseline is not a fact you can look up. It is a set of expectations you build by looking at the estate when nothing is wrong. Four are worth acquiring now, because exercises later depend on them.
WHERE SIGN-INS COME FROM 10.0.1.x is Manchester. Bristol has its own range. Anything else is a remote worker, a partner VPN egress, or an intrusion, and which one is a question about the ACCOUNT rather than about the address. WHEN SERVICE ACCOUNTS RUN 33 of 36 svc- logons go to SRV-NGE-MCR hosts, on schedule. The three that do not are an exercise later in this course. WHAT A NORMAL FAILURE RATE LOOKS LIKE A handful a day. People mistype passwords. "Some failures" tells you nothing; the brute force is identified by SHAPE, one address against one account 84 times. HOW MUCH NOISE POWERSHELL MAKES 17 hosts, constantly, legitimately. Any rule that starts "alert on PowerShell" dies on contact with this estate.
The baseline is what a model cannot have. Every fact in that card is a property of this estate. None of it is in any model's training data and no amount of prompting will surface it.
Where a generated answer goes wrong, it is very often because it applied a general truth to a specific estate where the general truth does not hold.
A note on how the corpus was built
The telemetry is generated rather than captured, and the way it was generated matters for what you can conclude from it.
Ordinary activity was simulated: users with roles, working patterns, devices and habits, producing the volume and rhythm a real estate produces. The attacks were then run through that activity as chains, so the malicious events sit inside the noise in the right proportion rather than being appended to it.
HOW THE CORPUS WAS ASSEMBLED
1 users given roles, hours, devices and habits
2 ordinary activity simulated from those, at real volume
3 the four attack chains RUN THROUGH that activity
NOT: attacks generated separately and appended. The order in
step 3 is why a wrong query returns ordinary rows rather than
returning nothing.
The practical consequence is that absence is meaningful here. If a query finds nothing, the reason is either that the thing did not happen or that you looked in the wrong place, and both are real answers. The corpus is not a puzzle where every query is supposed to return something interesting, and treating it as one will cost you a finding.
The same log line, twice
An accurate translation, and a findingHere is a single event, presented without comment.
Both readings are accurate. Only the second is a finding, and the difference is the thirty-five other logons that reader has seen.
2026-03-12T22:08:47Z svc-sql NE-LEWIS-LT Network Ntlm Success
Ask an assistant what it means and you will get an accurate reading: a service account named svc-sql authenticated to a host called NE-LEWIS-LT over the network using NTLM, and it succeeded. Every word of that is correct.
Now the same line read by someone who knows Northgate. svc-sql is a database service account. NE-LEWIS-LT is a laptop. Service accounts in this estate authenticate to SRV-NGE-MCR servers, thirty-three times out of thirty-six. NTLM rather than Kerberos is the older protocol and this estate uses it for legacy applications on servers, not for laptops. And svc-sql appears exactly once in thirty days of logon history, which is this line.
What "context" turns out to mean. People say a model lacks context as though context were a paragraph that could be supplied. In this example the load-bearing context is a distribution: what these accounts usually do, expressed as a ratio nobody has ever written down. It exists in the analyst's head as an expectation, and it got there by looking at the estate over months.
Test it on your own assistant
sign-ins would you expect on an ordinary weekday, and how many would indicate a password spray attack?
What this demonstrates. The answer is a plausible generalization, and watching one arrive is the useful part.
WHAT IT TOLD YOU a typical estate sees a handful of failed
sign-ins per user per day
WHERE THAT CAME FROM the question asked for a number, so a
number was produced
WHAT SETTLES IT one summarize over your own SigninLogs
YOUR ANSWER DEPENDS ON YOUR MFA POSTURE, YOUR LOCKOUT POLICY AND
HOW MANY SERVICE ACCOUNTS YOU RUN. NONE OF THAT WAS AVAILABLE.
Numbers arrive because numbers are what the question asked for, not because anything was measured, and your own estate settles it in one query.
Four kinds of context
Which can be supplied to an assistant and which cannotIt helps to separate them, because they differ in whether they can be supplied.
Stated facts. The estate has 810 staff, two offices, an M365 E5 tenant, a Palo Alto perimeter. Writeable, short, easily supplied.
Conventions. Hosts are named SITE-FUNCTION-NN. Service accounts are prefixed svc-. Office ranges are 10.0.1.x and Bristol's equivalent. Writeable with a little effort, and genuinely useful when supplied.
Distributions. How often service accounts touch laptops. What a normal daily failure count looks like. How many files finance downloads at month end. Not writeable, because they are not facts you hold as sentences. You hold them as expectations, and you discover you hold them when something violates one.
CAN YOU WRITE IT DOWN?
Stated facts yes, in one line 810 staff, E5, two sites
Conventions yes, with effort SRV- is a server
Distributions NO you hold these as
expectations, not sentences
Exceptions partly, never fully the list is generated by
events rather than design
THE TWO YOU CANNOT WRITE DOWN ARE THE TWO THE FINDINGS COME FROM.
Exceptions and history. That svc-legacyerp authenticates without MFA because of an application nobody has rewritten, that a partner's VPN egress looks foreign every Tuesday, that the alert on this host last month turned out to be a misconfigured backup agent. Partly writeable and never complete, because the list is generated by events rather than by design.
Two of the four fit in ten lines at the start of a session. The other two are why an assistant cannot reach a finding on its own.
The first two can be given to an assistant. The second two are where the findings are, and they resist being written down.
What you can actually supply
Ten lines of preamble, and what it will not fixSupplying the writeable half produces a real improvement, and three things carry most of it.
The naming conventions, because they let an assistant read a hostname as structured rather than as a string. Told that SRV- means server and NE-SURNAME-LT means a user laptop, the shape of the anomaly becomes visible from the names alone.
ONE LINE OF PREAMBLE, AND WHAT IT CHANGES
YOU WRITE "SRV- is a server. NE-SURNAME-LT is a user laptop."
BEFORE "svc-sql authenticated to NE-LEWIS-LT"
a hostname, treated as a string
AFTER "a service account authenticated to a USER LAPTOP"
a shape, and the shape is the anomaly
The address ranges, because internal and external is the most common distinction in security reasoning and it is entirely estate-specific.
The schema is the single highest-value thing to supply, because it attacks failure modes 1 and 4 directly. A model told which of the two logon tables holds domain authentication is much less likely to query the wrong one.
SAME REQUEST: "did this service account touch any laptops?"
WITHOUT THE PREAMBLE queried DeviceLogonEvents, guessed at what
a laptop is, matched on a name pattern it
invented
WITH IT queried IdentityLogonEvents, used NE-*-LT
because you said that is a user laptop
WHAT CHANGED the table and the convention. NOT whether
three logons to a laptop is normal here.
A reusable preamble is worth writing once. Ten lines describing conventions, ranges and the tables you use, pasted at the start of a session, measurably improves query generation and costs nothing after the first time. It does not make the assistant know your estate; it makes it stop guessing about the parts you could have told it.
The last line is the boundary. Ten lines of preamble fixed which table to query and how to read a hostname, and it did not touch the only question that decides whether three logons matter. That question is yours, and it stays yours no matter how much context you supply.
Running your first query
A baseline against the estate query to runConfirm the mechanism before an exercise depends on it. The block below is editable and runs against the live corpus.
SigninLogs
| where IPAddress == "45.83.64.117"
| summarize Attempts = count() by ResultType
| sort by Attempts desc
You should get three rows: 50126 with 58, 50053 with 25, and 0 with 1. Wrong password, account locked, and one success.
Every exercise in Module 3 is a variation on a query that returns something reasonable while excluding that row.
If the block did not run. The query surface needs a signed-in account with an active subscription. If you are reading the free preview, the exercise blocks in later modules will not execute, and the reasoning in each one still stands: every generated artifact in this course is accompanied by the finding it produces, so you can follow the argument without running anything. You will get considerably more out of it running the queries.
Practice
Build a baseline for your own environment hands onYou cannot borrow Northgate's baseline. The four expectations above are worth something only because somebody ran the queries that produced them, and the same four questions have different answers in your estate.
- How many sign-ins on an ordinary day?
- From how many distinct addresses?
- Which accounts authenticate most?
- What does the daily failure count look like?
- Which hosts run PowerShell, and how often?
Answer them against Northgate, then against your own estate, and write the answers somewhere you will find them. When a generated query returns a number, the only thing that tells you whether it is plausible is knowing what normal looks like, and normal is not something you can look up.
Next: section 0.5 turns the six modes into a prediction you can make from your own request, before you have read a single row of output.