Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
0.8 What This Course Builds
Introduction
This course is ten teaching modules and a Project, after this orientation, and each module adds a set of operators and the judgment to use them. This last sub of the orientation walks through what each module builds, with one short query from each part of the course.
The queries are deliberately ordinary, about working hours, mail and applications, so that the operators rather than any findings are the point. You'll derive a column and summarize by it, as the first modules teach, and summarize summaries, as Module 4 does.
You'll look at the text of mail as Module 6 would, and write a function with a parameter, as in Module 7. You'll build a time series that keeps every day, as Module 8 does, and measure a table before trusting it, as Module 10 does. Then you'll start the query library that the Project at the end completes.
By the end of the sub you should know what each module is for and what it assumes, where any particular question belongs, and what you will have built when the course is done.
The diagram is the course in one picture, read left to right and top to bottom. Ten modules, each building on the ones before it, and a Project that applies all of them to your own data. The first row is the foundations, the second the techniques that turn queries into analysis.
The Foundations
Modules 1 to 3The first three modules teach the work every query does, whatever its subject. Module 1 covers the structure of a query and the types of data it reads, extending what Sections 0.1 and 0.2 began.
Module 2 covers the ways to find rows: by value when the column is unknown, by term the way the engine indexes text, by time window, by address range, inside nested fields, and with negative filters and what they silently keep. Module 3 covers shaping results: choosing and computing columns, classifying values, ordering and limiting rows.
SigninLogs
| extend Hour = hourofday(TimeGenerated)
| summarize SignIns = count() by Hour
| top 3 by SignIns desc
Deriving an hour from each sign-in's time and grouping by it shows when the company works, in UTC: 14:00 UTC is the busiest hour, then 16:00 and 13:00. extend and summarize are the operators most queries in the course use, and these modules make them automatic.
The completed query below writes the derived column.
14:00, from a column that did not exist until the query made it, then summarized like any other. Deriving the value a question needs, rather than looking for a column that already holds it, is the first big step in writing queries.
The foundations take longer than they look, and that is deliberate. Most mistakes in advanced queries are foundation mistakes in disguise: a filter that matches too little, a type compared wrongly, a column with nulls treated as complete. Time spent on the first three modules is repaid in every later one.
Each foundation module also teaches the checks that go with its operators: in Module 1, how types affect comparisons; in Module 2, what negative filters silently keep; in Module 3, how ordering and limiting can hide the rows that matter. The checks are taught where the mistakes are made, rather than collected at the end.
Counting and Combining
Modules 4 and 5Modules 4 and 5 turn rows into answers, first within one table and then across several. Module 4 covers aggregation: counting correctly, describing a distribution, several questions in one pass, collecting values inside a group, rates per period, and the averages of averages that mislead. Module 5 covers combining tables: union for stacking, join in all its kinds, and the checks that keep a join honest.
SigninLogs
| summarize Apps = dcount(AppDisplayName) by UserPrincipalName
| summarize Accounts = count(), MedianApps = percentile(Apps, 50), MostApps = max(Apps)
Summarizing once per account, and then once more across accounts, two summaries in a row with nothing between them, shows that a typical account signs in to 12 applications and none to more than 13. A summary of summaries is a common shape in investigations, and Module 4 teaches it along with the ways it can mislead.
Module 5's joins are where many of the course's investigations first connect two kinds of evidence, from two tables: a sign-in with a device logon, an email with what its recipient did next. They are also where Section 10.4's multiplying rows come from, which is why Module 5 teaches the checks alongside the operators.
Together, Modules 4 and 5 are where most investigation queries live. A large share of the questions an analyst asks in a day are a filter, a summary and perhaps a join, and these two modules make those queries reliable as well as quick to write.
Combining tables is also where the course's month starts to make sense as a whole. Sign-ins, device logons, mail and directory changes each describe part of what happened; joined on accounts, devices and times, they describe it together. Many of the later modules' findings depend on a join first learned in Module 5.
The checks taught in these modules are also the ones most often skipped in practice: whether a distinct count is an estimate, whether a join kept every row it should, whether an average of averages means anything. Each is a one-line query, and each is taught where its mistake first becomes possible.
Messy Data and Reusable Logic
Modules 6 and 7Real data is messy, and real queries get reused; Modules 6 and 7 handle both. Module 6 covers text and nested data: parsing fields out of messages, key-value logs, regular expressions, decoding, JSON, property bags and arrays. Module 7 covers let for values and tables, functions and saved parsers, watchlists, external data, indicators and address ranges, which together turn queries into tools.
EmailEvents
| where EmailDirection == "Inbound"
| summarize Messages = count() by SenderFromDomain
| top 3 by Messages desc
Counting inbound mail by sending domain shows who writes to the company most. The domains are business services: Salesforce, Xero, Slack. Here the sender's domain is a ready-made column; much of Module 6 is about the many cases where it is not, and a value has to be pulled out of a message, a URL or a nested field before it can be counted.
The exercise below adds the filter the question needs.
Salesforce, Xero and Slack, once the company's own outbound and internal mail is set aside, which a filter on direction does in one line. Without it, the company's own domain, ne.com, tops the list with 2,501 messages and pushes Slack out of the top three. Choosing the rows a question is about is as much a part of Module 6's work as extracting values from them.
let FailedFor = (minimum: long) {
SigninLogs
| where ResultType != "0"
| summarize Failures = count() by UserPrincipalName
| where Failures >= minimum
};
FailedFor(10)
| count
A function with a parameter, FailedFor, writes the logic once, between braces, and calls it with a threshold. FailedFor(10) finds four accounts; any other threshold is one change away, and the function could be saved for every analyst to use. Module 7 builds functions like this, and the lookups and external lists that feed them.
Reusable logic changes how a team works. A function written once and shared means everyone counts failures, defines an unusual address or classifies a device the same way, and a correction made to the function corrects every query that calls it. Module 7 shows how to build and share such functions in a workspace.
Functions also make long queries readable. A query that calls FailedFor(10), then joins the result to the identity table, reads as two steps with names, where the same query written out in full would be a dozen lines whose purpose has to be worked out. Naming is the cheapest way to make a query explain itself.
Building in Practice
Record, routine and the common mistakesAcross the sub so far, four short queries, none longer than eight lines, showed what the first seven modules build: derived columns and summaries, summaries of summaries, values from text, and reusable logic.
The record collects what each part of the course builds.
What the course builds
By the endThe first five rows are the modules; the last row is the course's destination. Everything above it is in service of being able to write twenty good queries on data you know, check each one, and turn one into a rule.
The first three rungs are the module itself; the fourth is the test of it. The fourth rung is where learning consolidates. The challenge sets ask questions the subs did not, about the same month, and answering them without a sub to follow is the test of each module.
The second and third mistakes waste the course's design. The first mistake is tempting for readers who already write KQL. The later modules use the earlier ones constantly, and the habits taught early, checking joins, reading nulls correctly, filtering on time first, are what make the advanced techniques work.
Messy data is where many analysts stall. A log line with the needed value buried in free text, a JSON field that changes shape from one source to another, an array of steps where one matters: each blocks a query until it is unpacked. Module 6's operators make those values ordinary columns, after which every earlier technique applies to them.
Module 7 is also where a team's shared knowledge goes into queries: the list of the company's own addresses, the accounts that should never sign in from abroad, the indicators from a threat report. Kept outside the query in a watchlist or a file, such lists can be updated without touching the queries that use them.
Analysis
Modules 8 and 9Modules 8 and 9 are where queries stop retrieving and start analyzing, comparing each value with others. Module 8 covers time: series that include empty periods, comparing one period with another, one series per entity, window functions that compare each row with its neighbors, sessions and running totals. Module 9 covers relationships: graphs of accounts, addresses, applications and devices, and the patterns and paths that connect them.
EmailEvents
| make-series Messages = count() default = 0 on Timestamp
from datetime(2026-02-13) to datetime(2026-03-15) step 1d
| project Days = array_length(Messages),
Fewest = series_stats_dynamic(Messages).min, Most = series_stats_dynamic(Messages).max
A daily series of all mail, with every day of the month present, ranges from 22 messages on the quietest day to 776 on the busiest, a spread that a single monthly total would hide entirely. make-series guarantees that a day with no events appears as a zero rather than disappearing, which is what lets Module 8 find a day that breaks a pattern, or a quiet stretch that should not be quiet.
Module 9's graphs answer a different kind of question, one that joins can answer only with effort, a join per step of a path: not when something happened, but how things are connected. An address that reached several accounts, an application granted access by several people, a process tree from one program to everything it started. The course reaches graphs late because they combine almost everything before them.
These two modules are also where the month's attacks become clearest, because attacks are patterns over time and across relationships. A single sign-in rarely looks wrong; a sequence of them, or a connection between an address and several accounts, often does. Modules 8 and 9 give the tools to see such patterns in data of any size.
Module 9 introduces graphs from the beginning, on relationships the earlier modules have already found by other means, so that each graph query can be compared with a join that answers the same question.
Time also needs care that the earlier modules only begin to give it. A count per day is easy; a count per day that includes the days with nothing, compared with the same day last week, for each account separately, is Module 8's work, and it is what turns a pile of events into a pattern that can be read.
Trust
Module 10Module 10 is about keeping queries fast and answers right, in eight subs: how a query runs, why filtering early and choosing string operators carefully matter, how joins and summaries grow, the limits each portal enforces, and the checks that catch a borrowed, generated or careless query before its answer is relied on.
DeviceProcessEvents
| summarize Starts = count(), Devices = dcount(DeviceName), Programs = dcount(FileName)
Its first habit is knowing the size of what a query reads, before asking what it costs. The process table holds 9,576 starts of 39 programs on 18 devices, measured in one summarize; every question about what a query costs, and whether its answer is complete, starts from numbers like these.
Module 10 comes last because it applies to everything before it, and because its checks make the most sense once there is something substantial to check. Each technique in Modules 1 to 9 can be used well or carelessly, and Module 10 is where the course measures the difference.
Trust is also what makes the earlier modules usable at work. A query that answers a question in the lab is a start; a query whose cost is known, whose limits are respected and whose answer has been checked can be given to a colleague, put in a report or turned into a rule. Module 10 is the bridge between the two.
The modules' order is the order in which their techniques depend on each other: first finding things, then counting them, then connecting them, then seeing patterns, and finally making the whole reliable.
Module 10's last two subs take the checks furthest: adapting queries written by others for different data, and checking queries written by AI assistants. Both are common sources of KQL in practice, and both can produce queries that look right and are not. The same four checks, on names, values, grain and known cases, apply to both.
The Project
A library of your ownThe teaching modules end with a Project, which applies the whole course to data you know. It asks for twenty queries on your own data, each answering a stated question, each checked by the methods of Module 10, each with its table's pitfalls and its cost noted.
One of the twenty is taken further, to the quality of a scheduled rule: tested against a known case, with its threshold measured and its look-back chosen.
// Question: which applications do the most accounts use?
// Period: the sample month. Check: AllAccounts is every account that signed in.
let Total = toscalar(SigninLogs | summarize dcount(UserPrincipalName));
SigninLogs
| summarize Accounts = dcount(UserPrincipalName) by AppDisplayName
| top 3 by Accounts desc
| extend AllAccounts = Total
A finished entry looks like this, and every entry in the Project follows the same form: two comment lines with the question, the period and the check, a let for the total, and a query whose every row carries the figure it should be read against. SharePoint Online reached 94 of the month's 101 accounts, nearly all of them. Nothing about the entry needs explaining to the next reader.
The library is the course's lasting output, more than any score, because it keeps working, and growing, long after the course itself is finished. Its queries are written in your words, for questions you care about, on tables you know, and checked to the standard this course teaches. It is the thing you keep after the course is over, and it grows with every investigation afterwards.
Starting it now costs nothing, and the exercise below begins it with one complete entry of its own, ready to keep. Every sub's exercises produce queries; any that answer a question you might ask again can go into the library with a comment stating the question, and by the Project the work is mostly choosing and refining.
Choosing the twenty questions is part of the work, and worth doing early. A good library covers each family of tables an analyst uses, each rung of the ladder, and the questions that come up repeatedly in that analyst's work: who signed in from where, what ran on a device, what arrived by mail, what changed in the directory.
The three questions the practice task below asks you to write down are a first draft of that list.
Writing for the library also improves the queries themselves. A query that has to state its question in a comment is a query whose question has been thought about; one that has to state its check is one that has been checked. The discipline of the entry format is most of what Module 10 teaches, applied a query at a time.
The rule-quality query is the step from analysis to detection. It is held to the standard of something that runs unattended: a measured threshold, a look-back chosen for its schedule, a test against a case known to be in the data, and a note of what it costs. Building one is the course's last exercise and the first step toward detection engineering, which the platform teaches in its own course.
Start the Library
Your turnThe exercise writes the library's first entry: a query with its question as a comment, answering how many different accounts use each application.
The grader checks the columns, Accounts and AppDisplayName. Any order of applications is fine as long as the most-used come first.
The query is three lines and a comment. Read the entry as a template for the rest. A question in words, a period if the data is not fixed, a query, and, when it matters, a check. Twenty of these, chosen for your own environment, is what the course builds.
Practice
Read each module's introduction, work every sub, take the knowledge check, and do the challenge set, saving the queries worth keeping.
Module 0 ends here, with its Course Fit Check on the module's index page. Module 1 begins with query structure and data types.