In this section

0.1 What a Microsoft SOC Does

Module 0

Introduction

A security operations center, a SOC, is the team that watches an organization's systems for attacks and responds when it finds one. In a company that runs on Microsoft 365, the SOC works in Microsoft Defender and Microsoft Sentinel, and almost all of its work passes through one place: a queue of incidents, each grouping the alerts that describe one piece of suspicious activity. This course teaches the work done in that queue, from both of the seats in it: the analyst who investigates incidents, and the engineer who keeps the detections, automation and data behind them working. This first sub shows what that work looks like, using a month of real-shaped records from Northgate Engineering, the company every module in the course works with. Northgate is fictional, an 810-person engineering firm on Microsoft 365 E5, and its month of records was built for teaching: it contains the ordinary activity of a working company and several real attacks, and every figure in this sub comes from a query you can run.

A month in Northgate's SOC Signals sign-ins, mail, devices 710 alerts 39 kinds, 7 sources 540 incidents the SOC's queue 280 by Microsoft XDR resolved automatically 98 by analysts people in the SOC 59 by the SOC's automation an agent and a tuning rule 103 still open 102 never picked up The queue is where every part of the SOC's work meets.

The diagram is the whole course in one picture. Signals from the company's systems become alerts, alerts group into incidents, and incidents are closed by people, by automation, or by an AI agent, or they wait. Every module that follows teaches one part of that flow: how alerts are made, how incidents are investigated and closed, and how the whole thing is kept working and measured.

01

The Queue

Two hundred and two incidents in a month

Everything in a SOC eventually becomes an incident or doesn't. An incident is the unit of work: it has a status, an owner, a classification when it's closed, and a history of every change made to it. Microsoft Sentinel keeps that history in a table called SecurityIncident, and counting the latest state of each incident gives the month's queue.

SecurityIncident
| summarize arg_max(TimeGenerated, Status) by IncidentNumber
| summarize Incidents = count(), Closed = countif(Status == "Closed"),
    Open = countif(Status != "Closed"), NeverPickedUp = countif(Status == "New")

Five hundred and forty incidents arrived in the month, about eighteen a day, and 437 were closed by the end of it. For a company of 810 people that's an ordinary month rather than a quiet or a terrible one, which is what makes it useful for teaching: nothing in it is exaggerated to make a point. A hundred and three were still open, and 102 of those had never been picked up at all: their status was still New, no owner, no work recorded. Those four numbers already say most of what a SOC lead would want to know about the month, and every one of them comes from a single table you can query yourself in this course's lab. An incident carries more than a status, and the fields it records are the ones the whole course reads.

What an incident records

SecurityIncident

Status

New, Active or Closed; New means nobody has touched it

Owner

The person, or agent, it is assigned to

Classification

On closure: true positive, benign positive or false positive

Alerts

The alerts grouped into it, by their IDs

History

A new row for every change, with who made it and when

The last row is the one that surprises people, because most tables people meet hold one row per thing, and counting rows against incidents shows why it matters here.

SecurityIncident
| summarize Rows = count(), Incidents = dcount(IncidentNumber)

Twelve hundred and seventy-eight rows for 540 incidents, more than two per incident on average.

SecurityIncident doesn't hold one row per incident; it holds one row per change, so an incident that was created, assigned and closed has three rows. Taking the latest row for each incident, the arg_max step, is what turns a table of changes into a count of incidents. The same step appears in almost every query in this course that counts incidents, and it's the first thing to check when a count looks too high. Modules 11 and 13 come back to why getting that step wrong produces confident, wrong numbers. Three misreadings of an incident count are common enough to name now.

Reading an incident count
Rows are incidents.
1,278 rows describe 540 incidents; each change writes a row.Take the latest row per incident.
Closed means handled correctly.
Five of the month's closures without a person were found wrong.Closed means someone decided.
New means recent.
The oldest New incident had waited a month.New means untouched.
Every count in a SOC depends on reading the table the right way.

The middle row is the one the course returns to most. A status records what someone did, not whether it was right, and checking the second is most of what the later modules teach.

The 102 incidents nobody picked up are worth pausing on. They aren't a failure of effort in any one person; they're what happens when a queue receives more than its workers can read, and nobody has decided what to leave. Section 05 looks at what they were. For now, notice what the four numbers leave out: whether the 437 closures were right. A closed incident says someone decided; it doesn't say the decision was correct. Much of this course is about checking that, because a SOC that counts closures without checking them can look healthy while missing attacks.

02

Where the Alerts Come From

Seven sources, thirty-nine kinds

Incidents are built from alerts, and alerts come from products that watch different parts of the company. Counting the month's alerts by the product that raised them shows how varied the SOC's input is.

SecurityAlert
| summarize Alerts = count(), Kinds = dcount(AlertName) by ProviderName
| sort by Alerts desc

Seven sources raised 710 alerts of 39 different kinds. The names in the first column are the product codes Sentinel records, and they're worth translating once, because they appear throughout the course. Several are older names that stuck: MDATP is Defender for Endpoint, from the days when it was called Advanced Threat Protection, and OATP is Defender for Office 365 for the same historical reason. The codes don't change when the products are renamed, which is why a query written years ago still works, and why a new analyst needs the translation once.

Who raised the month's alerts

Sources

MCAS

Defender for Cloud Apps: activity in cloud applications

IPC

Entra ID Protection: risky sign-ins and users

OATP

Defender for Office 365: mail, including user reports

Defender for Identity

On-premises Active Directory

MDATP

Defender for Endpoint: devices

ASI Scheduled Alerts

Sentinel analytics rules the SOC wrote

The spread matters more than any single number. Cloud-application alerts were the most numerous, identity alerts the most varied, and the SOC's own analytics rules, the last row, raised only three. Most of what a Microsoft SOC sees is raised by Microsoft's products; the SOC's own detections fill the gaps those products leave, which is why Modules 2 to 6 spend so long on writing them well.

The source of an alert also decides where its evidence lives. An identity alert is investigated in the sign-in logs, a mail alert in the mail tables, a device alert in the device tables. Knowing which product raised an alert is the first step in knowing where to look. An incident can hold one alert or many, and Defender decides which alerts belong together.

SecurityIncident
| summarize arg_max(TimeGenerated, AlertIds) by IncidentNumber
| summarize Incidents = count() by Alerts = array_length(parse_json(AlertIds))
| sort by Alerts asc

Most incidents held a single alert, and about one in five held several, grouped because they came from the same product about the same user or device within a short window. The two largest held eight each; one of them is the domain compromise in Section 06, the most serious attack of the month, and even that incident left the attacker's cloud sign-in out. An analyst reads an incident's alerts, and then asks what else happened that the incident doesn't contain. That second question is where most investigations find their answer, and it's why the course teaches the tables under the incidents rather than the incident page alone.

03

Who Closes What

People, automation and an agent

A modern SOC isn't only people. Some incidents are closed by automation rules the engineers wrote, and some, increasingly, by AI agents. The incident table records who made the last change to each closed incident, which answers the question directly.

SecurityIncident
| summarize arg_max(TimeGenerated, Status, ModifiedBy) by IncidentNumber
| where Status == "Closed"
| summarize Incidents = count() by ClosedBy = case(
    ModifiedBy == "Phishing Triage Agent", "the agent",
    ModifiedBy startswith "Alert tuning", "automation",
    ModifiedBy endswith "@ne.com", "an analyst", ModifiedBy)

The largest group, 280, were resolved automatically by Microsoft XDR, the Defender platform itself: low-severity incidents, incidents whose alerts had all resolved, and identity risks users cleared themselves by completing multifactor authentication. Ninety-eight were closed by analysts, 51 by the Phishing Triage Agent, one of Microsoft's Security Copilot agents, and 8 by a tuning rule the SOC set up. Every one of those closures, whoever made it, is recorded with the same fields, so a person's work and software's can be compared directly. More than three quarters of the month's closures were made by software, and the course treats that as part of the job rather than a separate topic: an analyst increasingly checks what an agent or a rule concluded, as well as working incidents from scratch.

Software closes quickly and consistently, and it closes what it's given without knowing what it wasn't given. Module 13 found that three of the agent's closures were wrong, and that two of the tuning rule's 8 closed real phishing reports nobody ever looked at. Counted and checked, the same closures look different.

437 incidents closed
  98 by analysts
  339 by software
 
The month, as counted
✗
Says who decided.
437 incidents closed
  3 agent closures later reversed
  2 tuning-rule closures were phish
 
The month, as checked
✓
Says whether it was right.
Those errors were found by checking the records, which is the habit this course builds from its first module.

The split also changes what an analyst's day looks like. Where software closes the obvious cases, the analyst's queue holds the harder ones: incidents an agent kept open because they looked real, and alerts no rule could decide. The work gets more demanding per incident, and the checking of what software closed becomes part of it. Module 1 places the agent in the team as a worker to be supervised, and Module 11 measures it alongside the people. Neither module treats the agent as a threat to the analyst's job; both treat it as a colleague whose work, like anyone's, gets reviewed.

04

The Two Seats

The analyst and the engineer

The course is written for two roles that share one queue. They do different work on the same incidents, and most real SOCs mix them: a small team's analysts do engineering on quiet days, and its engineers work the queue on busy ones.

The two seats in the SOC

Northgate's month

The analyst

Works the queue: 98 incidents closed by people, and every true positive the agent kept open

The engineer

Keeps the queue working: the rules behind the alerts, the automation and the agent that closed 59

Shared

The same incidents, the same tables, and the same question: was this real, and what now?

The analyst's seat is the one most people picture: open an incident, work out what happened, decide whether it's real, contain it if it is, and write down what was found. The engineer's seat is less visible and shapes everything the analyst sees: the rules that raise alerts, the automation that closes the obvious ones, the agent that triages reported phish, and the data sources without which none of it works. When the engineer gets a rule wrong, the analyst drowns in false alarms or misses a real attack; when the analyst finds a gap, the engineer is who closes it.

The two seats also work in different parts of the same portal, which is worth knowing before the first module sends you there.

Where each seat works

Defender portal

Analyst

The incident queue, the incident page, advanced hunting

Engineer

Analytics rules, automation rules, data connectors, workspace health

Both

The same incidents and the same tables, read for different questions

Neither list is exhaustive, and the two overlap in advanced hunting, where both seats spend much of their time. Teaching both seats together is a deliberate choice. A SOC works when each seat understands the other's constraints, and the incident table above is the evidence: the same 540 incidents, handled by people, the platform, rules and an agent, all of which someone in one seat or the other is responsible for.

05

What Waits

Forty-four incidents nobody picked up

A queue always has a tail of work that isn't reached, and reading it tells a SOC more about itself than reading what was closed. The 102 incidents still at New can be grouped by the product whose alerts they contain, with the oldest of each.

let Waiting = SecurityIncident
    | summarize arg_max(TimeGenerated, Status, CreatedTime, AlertIds)
        by IncidentNumber
    | where Status == "New"
    | mv-expand Id = parse_json(AlertIds)
    | project IncidentNumber, CreatedTime, Id = tostring(Id);
SecurityAlert
| join kind=inner Waiting on $left.SystemAlertId == $right.Id
| summarize Waiting = dcount(IncidentNumber), Oldest = min(todatetime(CreatedTime))
    by ProviderName
| sort by Waiting desc

Sixty-three of the 102 came from cloud-application alerts. Identity alerts account for 33 more. The three from Defender for Endpoint are the ones that should worry a reader most: device alerts from 10 March onward, the week the ransomware reached a laptop, never opened. Nobody decided to leave these; they were simply never reached, and some of them, Module 13 found, were true positives in categories that did get worked.

The oldest waiting incident had been in the queue for a month, since the first morning the records cover. That doesn't mean it was dangerous; many cloud-application alerts describe routine activity. It means nobody knows, which is a different and worse position than knowing it's harmless, because an incident that was never read can't be counted as either a false alarm or an attack.

Faced with the tail, a new analyst usually proposes a rule for working through it. Make the call on the proposal below before reading where it resolves.

Judgment Call

Queue decision

Proposal, from a new analyst

"A hundred and two incidents are waiting, the oldest for a month. I'll work them oldest first, so nothing waits longer than it has to."

The 102 incidents at New: 63 from cloud-application alerts, the oldest from 13 February; 33 from identity alerts; 3 from Defender for Endpoint, from 10 March; and 3 others. Nobody has read any of them. Nobody has read any of them.

What is your call?

Pick one and write at least 15 words.

This is the tension every SOC lives with, and it doesn't go away with more staff; it only moves. The queue grows faster than people can read it, so something always waits, and the question is whether what waits was chosen or happened by default. Module 1 teaches how a SOC decides what to work first; Module 11 teaches how to measure what's waiting and for how long; Module 13 asks whether an agent could take some of it, and what it would need first.

06

What the Month Held

Six attacks in thirty days

The records aren't only a workload; they're a crime scene. Northgate's month contains several separate attacks, some of which connect, and the course investigates every one of them. In date order, the month's main events look like this.

What happened at Northgate in the month
27 Feba finance user's sessions taken over through a phishing relayidentity
2 Mara consent-phishing lure delivered, and an application reading a mailboxmail
9 to 12 Mara partner's mailbox taken over, sending phish to Northgatemail
11 to 14 Marthe on-premises domain compromised, to domain administratoridentity
12 Marransomware reaches a laptop and calls its command serverendpoint
15 Marthe compromised administrator signs in to the Azure portalcloud

Four of the six lines are marked because they're the ones that crossed from one part of the company to another. Attacks rarely stay in one product's view: a stolen password becomes a mailbox rule, a compromised laptop becomes a domain compromise, and each step raises its alert in a different place, if it raises one at all. A phishing relay that took over a finance user's sessions in February was used again in March to sign in as the administrator the domain compromise had captured, which makes two apparently separate incidents one intrusion. Ransomware reached a laptop and called out to its server. None of these was obvious from a single alert, and several were missed at first by the people, rules or agent that should have caught them.

Each attack is investigated in the module that owns its kind of evidence.

Where each attack is investigated

Modules

Phishing relay and the admin sign-in

Module 3, identity detections

Consent phishing and the partner mailbox

Module 4, mail detections

Ransomware on a laptop

Module 5, endpoint detections

The whole chain

Module 7, investigation playbooks

The split follows the evidence, not the attacker, because an analyst in the middle of an investigation needs to know how to read each kind of record, and each module teaches one kind thoroughly. By the end of the course, you'll have seen each attack from the alerts that caught it, the records that explain it, and the gaps that let parts of it through.

That's the real subject of the course: not the products, which Microsoft documents, but the work of turning 710 alerts into an accurate account of what happened, and keeping the operation that does it honest about what it gets wrong.

07

First Readings

Three assumptions the month contradicts

People new to security operations usually arrive with a picture of it, and Northgate's month contradicts three common parts of that picture.

Three first readings of a SOC
A SOC is people watching screens.
Three in four of the month's closures were made by software, and 102 incidents waited for anyone.It's people, rules, automation and agents on one queue.
The alerts are the work.
710 alerts became 540 incidents, and most of the work is deciding what each one means.The incidents are the work.
Engineering is separate from the queue.
Every rule, tuning rule and agent changes what the analysts see.The seats share the queue.
The queue is the SOC; everything else exists to make it workable.

The second row is the one that shapes how the course teaches. A SOC that tries to read every alert burns out its analysts on routine activity; one that reads incidents, and uses the alerts inside them as evidence, can keep up with a month like Northgate's. The difference between the two readings is stark set side by side.

710 alerts: the work
  every alert read
  every alert answered
 
A SOC drowning
✗
Nobody can read 710 alerts.
540 incidents: the work
  each a question about
  what the alerts mean
 
A SOC deciding
✓
The questions are the job.

The right-hand reading is the one the course builds. An alert is a claim that something happened; an incident is a question about what it means. Most of the analyst's time goes on that question, and most of the engineer's time goes on making sure the alerts that reach the question are worth asking about. The course gives both roughly equal weight for that reason. It also explains why the course keeps asking what a figure counts: 710 alerts sounds like a busy month, and 540 incidents sounds almost as busy, and the true workload is smaller than either, because the platform resolved more than half the incidents on its own, and it's set by how long each remaining incident takes to understand.

The third row is the reason the two seats are taught together, and Northgate's month shows the engineer's hand on the queue directly.

What the engineers changed this month

Changes to the queue

Tuning rule

Auto-Resolve on for reported phish until 17 February, then off

Agent

The Phishing Triage Agent triaging every report after that

Analytics rules

Three alerts from the SOC's own scheduled rules

Effect

Eight reports closed unseen before the change; every report triaged after it

Those changes were small, and each changed what the analysts saw. Turning off one tuning rule added every early report to the agent's work and to the analysts' review of it; turning on the agent took most reported phish off the analysts' queue altogether. An analyst who didn't know about either change would misread the month's numbers. A tuning rule closed two real phishing reports before any analyst saw them; an analyst's reversal of an agent's verdict is what showed the agent needed checking. Neither seat can do its job without knowing what the other did.

08

How the Course Follows the Queue

Thirteen modules, one flow

The course's modules follow the queue from the outside in: first how the SOC works, then where alerts come from, then what happens to an incident, then how the operation is kept working and measured.

How the course follows the queue
1How the SOC works
the queue, the team, the recordsModule 1
2What raises the alerts
detection methodology and the detection librariesModules 2 to 6
3What happens to an incident
investigation, documentation, response, automationModules 7, 8, 10
4What keeps it working
hardening, metrics, threat intelligence, AI agentsModules 9, 11 to 13
Each rung is a part of the same queue.

Each rung builds on the ones before it without depending on them: a module teaches its own subject completely, so a reader who arrives at Module 7 from a search can follow it. Read in order, though, the modules build one picture of Northgate's SOC, and the records you query in Module 1 are the same records the later modules investigate.

The rest of this orientation sets out the SOC's functions, where SOCs commonly fail, how their maturity can be measured, the pipeline an alert follows, what the course builds, and how to use its lab. None of it assumes experience in security operations; every term is explained where it first appears.

The course is built around records rather than screenshots because records are what a SOC actually works from. Portal screens change with every product update; the tables under them change far more slowly, and a query that answers a question today will answer it next year. Where the course shows a portal path, it's because the path is the only way to do something; everywhere else, it shows the query.

Practice

Read Northgate's queue yourself before reading anyone's summary of it.

Read a month of a SOC

In this course's lab; read-only.

  1. Run the queue query and write down the four numbers.
  2. Count the alerts by source, and note which source you would investigate in which tables.
  3. Count closures by who made them.
  4. List the incidents still at New, and pick the one you would open first, with a reason.

Keep your reason for the last step. Module 1 teaches how a SOC makes that choice, and your first answer is worth comparing with it.

The next sub, Section 0.2, sets out the four functions a SOC performs and where each shows up in Northgate's records.