Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
0.7 The Lab and How to Study
Introduction
Every module in this course is built around queries you can run, well over a hundred of them in all, each checked against the lab before publication: against Northgate's sample month in the course's practice console, and, once you're ready, against your own tenant. The queries aren't illustrations; they're how the course teaches, because the work it teaches is done by reading records, and the only way to learn to read records is to read them. This sub explains how the lab works, what the sample month contains, what's different about time in it, how to run the course's queries in your own environment without surprises, and the study habit that turns running queries into learning from them. It ends with a study record, a simple log that makes the habit stick.
The diagram is the habit, and it's the same for every query in every module. Every query in the course sits between two pieces of prose: the one above says what the query asks, and the one below says what the result means. Between them are three steps that are yours, and the most important is the one that's easiest to skip: writing down what you expect before you run anything. The dashed line is what happens when your prediction and the result disagree, which is exactly when the learning happens.
The Sample Month
What the lab holdsThe lab holds one month of records from Northgate Engineering, stored as the tables themselves, over two hundred thousand rows in all, the fictional company every module works with. It isn't a toy dataset: it's shaped like a real tenant's records, with ordinary activity from hundreds of people and devices, and several real attacks woven through it.
The sample month, in outline
Northgate EngineeringThe fourth row is the reason the month works for teaching. Most of what's in it is ordinary: people signing in, sending mail, running software. The attacks are a small fraction, as they are in a real tenant, which means finding them teaches the same skill a real SOC needs: separating the unusual from the merely uncommon. A dataset that was all attack would teach recognition; this one teaches search.
The month is also the same month in every module. An incident you meet in Module 1's queue is the same incident Module 7 investigates and Module 8 writes up, and the account whose sign-ins you read in Module 3 is the account whose mailbox Module 4 examines. Reading the course in order builds one picture of one company, and the modules refer to each other's findings freely.
Five of the main tables show the month's shape: how many rows each holds, and when its records start and end.
union withsource = Table SigninLogs, EmailEvents, DeviceProcessEvents,
CloudAppEvents, SecurityIncident
| extend Time = coalesce(TimeGenerated, Timestamp)
| summarize Rows = count(), First = min(Time), Last = max(Time) by Table
| sort by Rows desc
The sample month is also small enough to run anything in seconds, which a real tenant often isn't; a query that scans a month of mail in the lab in a second might take far longer against a real tenant's mail, so narrow the time window first when you run it at work. Each table covers the same thirty days, from the morning of 13 February to just before noon on 15 March. Mail is the largest, because every message is a row; incidents are the smallest, because they summarize everything else. The query itself shows something you'll need throughout the course, and the next section explains it.
Time in the Lab
A stopped clock, and two kinds of timestampTwo things about time in the lab catch almost everyone once. The first is the clock. The sample month ends at noon on 15 March, and the lab treats that moment as now, so any query that counts back from the present counts back from there.
SigninLogs
| where TimeGenerated > ago(1d)
| summarize SignIns = count(), Earliest = min(TimeGenerated),
Latest = max(TimeGenerated)
A hundred and thirty-six sign-ins in the last day, and the earliest is at noon on 14 March, exactly one day before the lab's now. In your own tenant, the same query counts back from the real present, which is why the course usually uses fixed dates in its examples: a fixed date gives the same answer in the lab next year, while ago() gives whatever the last day happened to contain.
The second is the timestamp column. The query in Section 01 needed a coalesce() because the tables don't all name their time column the same way: tables that come from Microsoft Sentinel's own collection use TimeGenerated, while Defender XDR's advanced hunting tables use Timestamp. It's a real difference in Microsoft's products, not a quirk of the lab. In the lab, a query that filters a Defender table on TimeGenerated returns nothing, with no error to say why.
Three readings of time in the lab are common enough to name.
The second row is the one that costs the most time when it's missed, in the lab and at work alike, because an empty result looks like an absence of activity rather than a wrong column. The course's queries always use the right column for each table, and when you write your own, checking the column is the first thing to do when a query returns nothing.
Running a Query
From the page to the consoleEvery query you can run carries a link above it, Open in Advanced hunting, which opens the course's practice console in a new tab. The console looks and behaves like Defender's advanced hunting, and it runs against the sample month. The query itself travels by copy and paste: use the Copy button on the query, paste it into the console, and run it.
The console is a practice version of the real thing, built for this course and kept deliberately close to it.
The practice console
What it isIf you have your own tenant open as well, keep the two in separate tabs, and check which tab you're in before running anything: the course's queries are safe anywhere, but a result read in the wrong tab is a result misread. The copy step is deliberate. Typing or pasting the query yourself, rather than having it run in place, puts it in front of you in the tool you'll use at work, where you can edit it, break it and fix it. Most of the learning in a query comes from changing it: widening a window, removing a filter, grouping by a different column, and seeing what changes.
Before you run anything, though, write down what you expect, even roughly: a number, a range, which row will be largest. A rough guess is enough; the point isn't to be right, it's to have something the result can contradict. Here's a simple query to practice on: every sign-in in the month, grouped by whether it succeeded.
SigninLogs
| summarize SignIns = count() by Result = iff(ResultType == "0", "success", "failure")
Seven thousand eight hundred and sixty-five successes and 338 failures, about four percent failed. The query treats any result code other than 0 as a failure, which is the usual shortcut: ResultType is a code, and 0 means success. If you predicted mostly success with a few percent failures, the result confirms a sensible model of how people sign in. If you predicted something very different, the difference is the lesson: either your model of a normal tenant needs adjusting, or you've learned that this month is unusual, and either is worth knowing before you start reading attacks in it.
Four percent is a figure worth remembering, because it's a baseline: when a later module shows an account with half its sign-ins failing, the contrast with four percent is what makes it stand out. The difference between running with and without a prediction is larger than it looks.
Read the query
run it
read the output
Nothing written down firstRead the query
predict: "mostly success,
a few percent failure"
run it, compare
A guess testedThe left-hand pane is how most people read a course's queries, and it has a hidden cost: after the fact, almost every result looks reasonable, so nothing surprises you and nothing is learned. The right-hand pane turns every query into a small test of your understanding. It takes ten seconds per query, and it's the single habit that most separates people who finish the course able to do the work from people who finish it having read about it.
Reading the Printed Output
What the page shows, and what it doesn'tBelow every runnable query, the course prints the output the lab returns, so you can check your result and read the explanation even when you're not at a console. The printed outputs follow a few conventions that are worth knowing, because they differ slightly from what the console shows.
Reading the printed outputs
ConventionsEach convention exists to make the page readable on a laptop or a phone, where a console's full output wouldn't fit. The first and third rows cause most of the small mismatches readers notice. The console shows full timestamps with seconds and time zone, where the page shortens them; and where a query doesn't sort, the console and the page may list the same rows in a different order. Neither is a real difference. A real difference is a different number, a missing row or an extra one, and the rest of this sub is about what to do when you see one.
When your result and the printed output differ only in order or in how a time is shown, move on; when a number differs, stop and use Section 07. Three readings of a printed output produce most false alarms.
The printed outputs are generated from the same lab the console runs, and every one is checked against it before the course is published. A query whose output didn't match its printed version would be a bug in the course, and worth reporting.
The prose after each output is where the course draws its conclusion: what the numbers mean, which one matters, and what it implies. Read it after you've compared the output with your prediction, not before, so that your own reading comes first and the course's second.
Your Own Tenant
Running the course's queries on your recordsEvery query in the course reads standard Microsoft tables, so it can be run in your own tenant's advanced hunting, read-only, once you have access. Running the course's queries on your own records is the best practice there is, because the attacks you'll meet at work are in your records, not Northgate's. It needs a little care the first time, and the same gate applies to every query.
names accounts, devices, domains in the examples
time datetime() values and ago() windows
tables the ones your workspace actually holdsChange those three, and nothing else, the first time.The gate is the same for every query in every module, which makes it a habit rather than a checklist to look up each time. Most of the changes are small, and all of them are in the query's text. Northgate's accounts end in ne.com, and its device names follow Northgate's scheme; replace them with your own. Fixed datetime() values point at Northgate's month; change them to a window in yours. And if a query reads a table your workspace doesn't hold, note the gap rather than forcing the query: a missing table is itself a finding, and Section 0.6's gate is a quick way to check them all at once.
It helps to know in advance which parts of a query travel and which don't.
What carries over to your tenant, and what doesn't
Lab and productionThe last row is the one to check first in a new workspace. A table you don't hold returns an error; a table your workspace keeps for thirty days can't answer a question about last quarter. Read-only queries can't change anything in your tenant; they only read. The care is in the reading: your month will have its own attacks, its own noise and its own gaps, and the course's prose explains Northgate's results, not yours. The skill the course is building is exactly the one you'll use here: reading a result you've never seen before and working out what it means.
How to Study
An order, and a paceThe course can be read straight through, and many readers will. It can also be read by need, once the foundations are in place, and for readers who work in a SOC already that's often the better route.
The later rungs are flexible because the modules were written to stand alone: each defines what it uses and points back to where an idea was first taught, so arriving at Module 5 from a search works. The first rung isn't optional, and not only because this sub says so. Module 1's vocabulary and Module 2's method are used everywhere after them, and a reader who skips them spends the detection modules working out what a specification or a triage priority is, instead of what the detection does.
Readers who already work in a SOC often find the later modules faster going than the early ones, because they recognize the attacks; the early modules are where the course's own habits are set, and those habits are what the later modules assume. Each part of the course ends the same way, which makes the rhythm predictable.
Pace matters less than rhythm, and rhythm is mostly about the queries. A sub is designed to be read in one sitting with its queries run, and a module in a week or two alongside other work. Reading several subs in a sitting without running the queries is faster and teaches much less; running every query once, with a prediction, is the rhythm the course is built for. Each sub ends with a practice card, a short list of things to do with what it taught, and each module ends with a knowledge check of eight scenarios that test whether the ideas carry to situations the sub didn't describe.
The knowledge checks are worth taking seriously even though they aren't graded. Each scenario is built so that the wrong answers are the ones the module's own misreadings would produce, which means a wrong answer points at exactly which misreading you still hold.
When a Result Doesn't Match
Five causes, and where to look firstSooner or later a query you run won't return what the page shows, or what you predicted. In the lab that's usually a copying problem; in your own tenant it's usually one of a handful of things, and checking them in order finds the cause quickly.
When a result doesn't match
Five causes, in order of likelihoodTime and grain cause most mismatches, in the lab and at work, which is why they're first. A window that's shorter than you thought, an ago() in a lab whose clock stopped, or a Timestamp column filtered as TimeGenerated all return fewer rows than expected, often none. Grain is subtler: the incident table writes a row for every change, so counting rows counts changes, and every count of incidents in the course takes each incident's latest row first. Module 11 and Module 13 both return to this, because it produces confident, wrong numbers.
Values catch people in their own tenants. A status written Closed in one table and closed in another, a domain written with and without a subdomain, or a user named by display name in one column and by address in another will all make a filter match nothing. A quick distinct on the column you're filtering, before the filter, shows how the values are actually written. On the incident table it's one line.
SecurityIncident
| distinct Status
Three values, each capitalized, and only three: New, Active and Closed. A filter written Status == "closed" matches nothing, because == is case-sensitive in KQL; =~ is the version that ignores case. The safe habit is to copy the value from a distinct rather than type it from memory.
These are the same checks Module 13 applies to queries Copilot writes, and Module 11 to the SOC's own measures, and for the same reason: a query that runs isn't a query that's right, whoever wrote it.
A Study Record
One line per queryThe last habit is the simplest, and the one most people skip after the first week. Keep a record, one line per query you run, with what you predicted and what you learned. It takes seconds per line and it does three things: it makes the prediction step unskippable, it gives you somewhere to put the lessons that mismatches teach, and by the end of the course it's a map of what you found hard.
A study record, one line per query
Keep it as you goThe record doesn't need a tool; a text file or a notebook page works, and so does a spreadsheet if you like sorting it. The last column is the one that matters. Most lines will say the prediction was right, and that's fine; the lines where it wasn't are your personal syllabus, the places where your model of how Microsoft's records behave was wrong and got corrected. Read them back before starting the Project, and you'll know which modules to revisit.
The record is also the first thing to reach for when the course sends you to your own tenant, because every line in it is a query you've already run once. A prediction written for Northgate and then tested against your records is a small experiment in how your environment differs, and the differences are exactly what a new analyst in a SOC spends their first months learning. You'll have started that before you arrive.
Practice
Start your study record now, with this sub's own queries.
Your first three lines
In the practice console; read-only.
- Predict, then run, the table ranges query, and note which table is largest.
- Predict, then run, the last-day sign-ins, and write down the lab's now.
- Predict, then run, the sign-in results, and record how close you were.
- Change one thing in each query, and predict again before running it.
Keep the record open for Module 1; the first query there is the queue from Section 0.1, and your prediction for it can use what this orientation has already shown you.
This is the last sub of the orientation. Module 1, How a Microsoft SOC Works, is next, and it opens with the queue.