In this section

0.7 The Lab and How to Study

Module 0

Introduction

Every module in this course is built around queries you can run, well over a hundred of them in all, each checked against the lab before publication: against Northgate's sample month in the course's practice console, and, once you're ready, against your own tenant. The queries aren't illustrations; they're how the course teaches, because the work it teaches is done by reading records, and the only way to learn to read records is to read them. This sub explains how the lab works, what the sample month contains, what's different about time in it, how to run the course's queries in your own environment without surprises, and the study habit that turns running queries into learning from them. It ends with a study record, a simple log that makes the habit stick.

The study loop, around every query Read what it asks Predict write it down Run in the console Compare three results Explain what it means a mismatch sends you back to read again The prediction is the step that does the learning.

The diagram is the habit, and it's the same for every query in every module. Every query in the course sits between two pieces of prose: the one above says what the query asks, and the one below says what the result means. Between them are three steps that are yours, and the most important is the one that's easiest to skip: writing down what you expect before you run anything. The dashed line is what happens when your prediction and the result disagree, which is exactly when the learning happens.

01

The Sample Month

What the lab holds

The lab holds one month of records from Northgate Engineering, stored as the tables themselves, over two hundred thousand rows in all, the fictional company every module works with. It isn't a toy dataset: it's shaped like a real tenant's records, with ordinary activity from hundreds of people and devices, and several real attacks woven through it.

The sample month, in outline

Northgate Engineering

Company

810 staff, Microsoft 365 E5, Entra ID, Defender XDR and Sentinel

Period

13 February to noon on 15 March 2026

Tables

50 tables and 8 watchlists, the ones a Microsoft SOC reads

Contents

A working company's ordinary activity, and the attacks the course investigates

The fourth row is the reason the month works for teaching. Most of what's in it is ordinary: people signing in, sending mail, running software. The attacks are a small fraction, as they are in a real tenant, which means finding them teaches the same skill a real SOC needs: separating the unusual from the merely uncommon. A dataset that was all attack would teach recognition; this one teaches search.

The month is also the same month in every module. An incident you meet in Module 1's queue is the same incident Module 7 investigates and Module 8 writes up, and the account whose sign-ins you read in Module 3 is the account whose mailbox Module 4 examines. Reading the course in order builds one picture of one company, and the modules refer to each other's findings freely.

Five of the main tables show the month's shape: how many rows each holds, and when its records start and end.

union withsource = Table SigninLogs, EmailEvents, DeviceProcessEvents,
    CloudAppEvents, SecurityIncident
| extend Time = coalesce(TimeGenerated, Timestamp)
| summarize Rows = count(), First = min(Time), Last = max(Time) by Table
| sort by Rows desc

The sample month is also small enough to run anything in seconds, which a real tenant often isn't; a query that scans a month of mail in the lab in a second might take far longer against a real tenant's mail, so narrow the time window first when you run it at work. Each table covers the same thirty days, from the morning of 13 February to just before noon on 15 March. Mail is the largest, because every message is a row; incidents are the smallest, because they summarize everything else. The query itself shows something you'll need throughout the course, and the next section explains it.

02

Time in the Lab

A stopped clock, and two kinds of timestamp

Two things about time in the lab catch almost everyone once. The first is the clock. The sample month ends at noon on 15 March, and the lab treats that moment as now, so any query that counts back from the present counts back from there.

SigninLogs
| where TimeGenerated > ago(1d)
| summarize SignIns = count(), Earliest = min(TimeGenerated),
    Latest = max(TimeGenerated)

A hundred and thirty-six sign-ins in the last day, and the earliest is at noon on 14 March, exactly one day before the lab's now. In your own tenant, the same query counts back from the real present, which is why the course usually uses fixed dates in its examples: a fixed date gives the same answer in the lab next year, while ago() gives whatever the last day happened to contain.

The second is the timestamp column. The query in Section 01 needed a coalesce() because the tables don't all name their time column the same way: tables that come from Microsoft Sentinel's own collection use TimeGenerated, while Defender XDR's advanced hunting tables use Timestamp. It's a real difference in Microsoft's products, not a quirk of the lab. In the lab, a query that filters a Defender table on TimeGenerated returns nothing, with no error to say why.

Three readings of time in the lab are common enough to name.

Reading time in the lab
ago(1d) means the last day before today.
The lab's clock stops at noon on 15 March; ago() counts back from there.Now is 15 Mar 12:00.
Every table has a TimeGenerated column.
Defender XDR tables use Timestamp; Sentinel tables use TimeGenerated.Check which, or coalesce them.
A query copied to your tenant returns the same rows.
Your tenant has your own month; the shape carries, the rows don't.Expect your own numbers.
Time is the most common reason a result surprises.

The second row is the one that costs the most time when it's missed, in the lab and at work alike, because an empty result looks like an absence of activity rather than a wrong column. The course's queries always use the right column for each table, and when you write your own, checking the column is the first thing to do when a query returns nothing.

03

Running a Query

From the page to the console

Every query you can run carries a link above it, Open in Advanced hunting, which opens the course's practice console in a new tab. The console looks and behaves like Defender's advanced hunting, and it runs against the sample month. The query itself travels by copy and paste: use the Copy button on the query, paste it into the console, and run it.

The console is a practice version of the real thing, built for this course and kept deliberately close to it.

The practice console

What it is

Language

KQL, as in Defender advanced hunting

Data

The sample month: the same tables every module reads

Clock

Now is noon on 15 March 2026

Reached from

The Open in Advanced hunting link on any runnable query

If you have your own tenant open as well, keep the two in separate tabs, and check which tab you're in before running anything: the course's queries are safe anywhere, but a result read in the wrong tab is a result misread. The copy step is deliberate. Typing or pasting the query yourself, rather than having it run in place, puts it in front of you in the tool you'll use at work, where you can edit it, break it and fix it. Most of the learning in a query comes from changing it: widening a window, removing a filter, grouping by a different column, and seeing what changes.

Before you run anything, though, write down what you expect, even roughly: a number, a range, which row will be largest. A rough guess is enough; the point isn't to be right, it's to have something the result can contradict. Here's a simple query to practice on: every sign-in in the month, grouped by whether it succeeded.

SigninLogs
| summarize SignIns = count() by Result = iff(ResultType == "0", "success", "failure")

Seven thousand eight hundred and sixty-five successes and 338 failures, about four percent failed. The query treats any result code other than 0 as a failure, which is the usual shortcut: ResultType is a code, and 0 means success. If you predicted mostly success with a few percent failures, the result confirms a sensible model of how people sign in. If you predicted something very different, the difference is the lesson: either your model of a normal tenant needs adjusting, or you've learned that this month is unusual, and either is worth knowing before you start reading attacks in it.

Four percent is a figure worth remembering, because it's a baseline: when a later module shows an account with half its sign-ins failing, the contrast with four percent is what makes it stand out. The difference between running with and without a prediction is larger than it looks.

Read the query
  run it
  read the output
 
Nothing written down first
✗
Every result looks reasonable.
Read the query
  predict: "mostly success,
  a few percent failure"
  run it, compare
 
A guess tested
✓
A surprise is a lesson.

The left-hand pane is how most people read a course's queries, and it has a hidden cost: after the fact, almost every result looks reasonable, so nothing surprises you and nothing is learned. The right-hand pane turns every query into a small test of your understanding. It takes ten seconds per query, and it's the single habit that most separates people who finish the course able to do the work from people who finish it having read about it.

04

Reading the Printed Output

What the page shows, and what it doesn't

Below every runnable query, the course prints the output the lab returns, so you can check your result and read the explanation even when you're not at a console. The printed outputs follow a few conventions that are worth knowing, because they differ slightly from what the console shows.

Reading the printed outputs

Conventions

Times

Shortened to day, month and time; the console shows full timestamps

Columns

Aligned for reading; long values wrapped onto a second line

Order

As the engine returns it; sort in your query if order matters

Empty cells

Blank in the data, not missing from the output

Each convention exists to make the page readable on a laptop or a phone, where a console's full output wouldn't fit. The first and third rows cause most of the small mismatches readers notice. The console shows full timestamps with seconds and time zone, where the page shortens them; and where a query doesn't sort, the console and the page may list the same rows in a different order. Neither is a real difference. A real difference is a different number, a missing row or an extra one, and the rest of this sub is about what to do when you see one.

When your result and the printed output differ only in order or in how a time is shown, move on; when a number differs, stop and use Section 07. Three readings of a printed output produce most false alarms.

Reading a printed output
A different order means a different result.
Unsorted rows can come back in any order.Compare the rows, not their order.
A shortened time means a different time.
The page trims seconds and zone for reading.Compare to the minute.
The output is the answer.
It is the evidence; the prose below says what it answers.Read both.
Most apparent mismatches are presentation; a real one is a different number.

The printed outputs are generated from the same lab the console runs, and every one is checked against it before the course is published. A query whose output didn't match its printed version would be a bug in the course, and worth reporting.

The prose after each output is where the course draws its conclusion: what the numbers mean, which one matters, and what it implies. Read it after you've compared the output with your prediction, not before, so that your own reading comes first and the course's second.

05

Your Own Tenant

Running the course's queries on your records

Every query in the course reads standard Microsoft tables, so it can be run in your own tenant's advanced hunting, read-only, once you have access. Running the course's queries on your own records is the best practice there is, because the attacks you'll meet at work are in your records, not Northgate's. It needs a little care the first time, and the same gate applies to every query.

A query, run in your own tenantCopied from the course, run read-only.Your data, your month.Before you trust the result
The gateHave you changed what is specific to Northgate?
names       accounts, devices, domains in the examples
time        datetime() values and ago() windows
tables      the ones your workspace actually holds
Change those three, and nothing else, the first time.
Yes: read the result as yoursThe shape of the answer should match the course; the numbers won't.
No: expect nothing backA query naming ne.com in your tenant returns nothing, and teaches nothing.
Read-only queries are safe; the care is in reading the answer.

The gate is the same for every query in every module, which makes it a habit rather than a checklist to look up each time. Most of the changes are small, and all of them are in the query's text. Northgate's accounts end in ne.com, and its device names follow Northgate's scheme; replace them with your own. Fixed datetime() values point at Northgate's month; change them to a window in yours. And if a query reads a table your workspace doesn't hold, note the gap rather than forcing the query: a missing table is itself a finding, and Section 0.6's gate is a quick way to check them all at once.

It helps to know in advance which parts of a query travel and which don't.

What carries over to your tenant, and what doesn't

Lab and production

Carries over

Table names, column names, operators, the shape of each answer

Changes

Names, domains, addresses, the dates, the volumes

Differs in kind

Which tables you hold, and how long your workspace keeps them

The last row is the one to check first in a new workspace. A table you don't hold returns an error; a table your workspace keeps for thirty days can't answer a question about last quarter. Read-only queries can't change anything in your tenant; they only read. The care is in the reading: your month will have its own attacks, its own noise and its own gaps, and the course's prose explains Northgate's results, not yours. The skill the course is building is exactly the one you'll use here: reading a result you've never seen before and working out what it means.

06

How to Study

An order, and a pace

The course can be read straight through, and many readers will. It can also be read by need, once the foundations are in place, and for readers who work in a SOC already that's often the better route.

A study order that works
1Module 0, then Modules 1 and 2
the vocabulary, the queue, and the methodin order
2The detection module closest to your work
identity, mail, endpoint or cloud firstby need
3Modules 7 and 8
investigation and reporting, using the detectionsin order
4The rest, as your work needs them
response, measures, intelligence, AIby need
Start in order; branch once the foundations are in place.

The later rungs are flexible because the modules were written to stand alone: each defines what it uses and points back to where an idea was first taught, so arriving at Module 5 from a search works. The first rung isn't optional, and not only because this sub says so. Module 1's vocabulary and Module 2's method are used everywhere after them, and a reader who skips them spends the detection modules working out what a specification or a triage priority is, instead of what the detection does.

Readers who already work in a SOC often find the later modules faster going than the early ones, because they recognize the attacks; the early modules are where the course's own habits are set, and those habits are what the later modules assume. Each part of the course ends the same way, which makes the rhythm predictable.

How each part of the course ends
1Each sub
a practice card: a few things to do with what it taughtevery sitting
2Each module
a summary, and a knowledge check of eight scenariosevery week or two
3The course
reference modules to return to, and the Projectat the end
The same shape every time, so the rhythm is predictable.

Pace matters less than rhythm, and rhythm is mostly about the queries. A sub is designed to be read in one sitting with its queries run, and a module in a week or two alongside other work. Reading several subs in a sitting without running the queries is faster and teaches much less; running every query once, with a prediction, is the rhythm the course is built for. Each sub ends with a practice card, a short list of things to do with what it taught, and each module ends with a knowledge check of eight scenarios that test whether the ideas carry to situations the sub didn't describe.

The knowledge checks are worth taking seriously even though they aren't graded. Each scenario is built so that the wrong answers are the ones the module's own misreadings would produce, which means a wrong answer points at exactly which misreading you still hold.

07

When a Result Doesn't Match

Five causes, and where to look first

Sooner or later a query you run won't return what the page shows, or what you predicted. In the lab that's usually a copying problem; in your own tenant it's usually one of a handful of things, and checking them in order finds the cause quickly.

When a result doesn't match

Five causes, in order of likelihood

Time

A window, an ago(), a Timestamp where TimeGenerated was expected

Grain

Rows per change counted as things; take the latest row first

Values

A status, domain or name spelled differently from the data

Tables

A table missing or empty in your workspace

The query

A copying error: a lost line, a changed quote

Time and grain cause most mismatches, in the lab and at work, which is why they're first. A window that's shorter than you thought, an ago() in a lab whose clock stopped, or a Timestamp column filtered as TimeGenerated all return fewer rows than expected, often none. Grain is subtler: the incident table writes a row for every change, so counting rows counts changes, and every count of incidents in the course takes each incident's latest row first. Module 11 and Module 13 both return to this, because it produces confident, wrong numbers.

Values catch people in their own tenants. A status written Closed in one table and closed in another, a domain written with and without a subdomain, or a user named by display name in one column and by address in another will all make a filter match nothing. A quick distinct on the column you're filtering, before the filter, shows how the values are actually written. On the incident table it's one line.

SecurityIncident
| distinct Status

Three values, each capitalized, and only three: New, Active and Closed. A filter written Status == "closed" matches nothing, because == is case-sensitive in KQL; =~ is the version that ignores case. The safe habit is to copy the value from a distinct rather than type it from memory.

These are the same checks Module 13 applies to queries Copilot writes, and Module 11 to the SOC's own measures, and for the same reason: a query that runs isn't a query that's right, whoever wrote it.

08

A Study Record

One line per query

The last habit is the simplest, and the one most people skip after the first week. Keep a record, one line per query you run, with what you predicted and what you learned. It takes seconds per line and it does three things: it makes the prediction step unskippable, it gives you somewhere to put the lessons that mismatches teach, and by the end of the course it's a map of what you found hard.

A study record, one line per query

Keep it as you go

Query

Module, section, what it asked

Prediction

What you expected, before running it

Result

What came back, and any mismatch with the printed output

Lesson

One sentence: what the difference taught you

The record doesn't need a tool; a text file or a notebook page works, and so does a spreadsheet if you like sorting it. The last column is the one that matters. Most lines will say the prediction was right, and that's fine; the lines where it wasn't are your personal syllabus, the places where your model of how Microsoft's records behave was wrong and got corrected. Read them back before starting the Project, and you'll know which modules to revisit.

The record is also the first thing to reach for when the course sends you to your own tenant, because every line in it is a query you've already run once. A prediction written for Northgate and then tested against your records is a small experiment in how your environment differs, and the differences are exactly what a new analyst in a SOC spends their first months learning. You'll have started that before you arrive.

Practice

Start your study record now, with this sub's own queries.

Your first three lines

In the practice console; read-only.

  1. Predict, then run, the table ranges query, and note which table is largest.
  2. Predict, then run, the last-day sign-ins, and write down the lab's now.
  3. Predict, then run, the sign-in results, and record how close you were.
  4. Change one thing in each query, and predict again before running it.

Keep the record open for Module 1; the first query there is the queue from Section 0.1, and your prediction for it can use what this orientation has already shown you.

This is the last sub of the orientation. Module 1, How a Microsoft SOC Works, is next, and it opens with the queue.