Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
0.2 The Four Functions
Introduction
Every SOC, whatever its size or tools, performs four functions. It detects possible attacks, it triages and investigates what it detects, it responds to what turns out to be real, and it improves all three using what it learns. The names vary from one organization to the next, and small SOCs have one person doing all four, but the functions don't change, and neither does the order in which an incident passes through the first three. This sub sets them out one at a time, and for each one it shows where the function leaves its mark in Northgate's records, because a function that leaves no record can't be checked, and a function that can't be checked is only hoped for. Nothing in this sub needs prior experience: each function is explained from what it does to Northgate's month, and each record is introduced where it first appears.
The diagram draws the four as a loop rather than a line. Detection, triage and response run left to right with every incident; improvement runs back, from what the SOC learned to how it detects next time. Most SOCs draw only the first three and wonder why the same problems return every month. The loop is also why the course's later modules keep returning to records the earlier ones produced: a verdict recorded in Module 1's queue is an input to Module 11's measures, and a sighting from Module 12 is an input to the detections of Modules 3 to 6. The box at the far right is the course's constant theme: every step writes a record, and the records are what the rest of this sub reads.
Four Functions
One month, four kinds of workThe four functions are easiest to understand side by side, with the evidence each left in Northgate's month. The figures below come from the queries in the rest of this sub and from Section 0.1.
The four functions at Northgate
One monthThe first three rows describe work that happens to every incident, in order, and at Northgate each one is done partly by people and partly by software. The fourth is different in kind: it doesn't happen to an incident at all, but to the SOC itself, and its records exist whether anyone reads them or not. That last clause is the important one, and it's why the fourth row reads read or not. A SOC's health records, cost figures and threat intelligence accumulate every day; improvement is the function that reads them and acts, and it's the one most often skipped when the queue is busy.
Ownership of each function sits mainly with one of the two seats from Section 0.1, though never entirely.
Which seat owns which function
OwnershipThe split is a starting point, not a boundary, and every real SOC blurs it somewhere, often daily. Engineers own most of detection and improvement; analysts own most of triage, investigation and response. Each seat needs the other's function to work: an analyst can't triage well on noisy detections, and an engineer can't tune detections without the analysts' verdicts. In a small SOC one person may hold both seats on the same day, which makes the split more important to state, not less: knowing which hat you're wearing tells you which records you owe the other seat.
Detect
Rules, products and the data under themDetection is the function that turns activity into alerts. At Northgate most detection is done by Microsoft's products, and a small part by analytics rules the SOC wrote in Sentinel. Both depend on data arriving: a rule that reads a table that stopped filling raises nothing, and raises nothing silently. Sentinel records the health of the pieces that do this work in a table called SentinelHealth.
SentinelHealth
| summarize Runs = count(), Failures = countif(Status != "Success"),
Resources = dcount(SentinelResourceName) by SentinelResourceType
Four kinds of component, each with its own row, and three of them had failures in the month. Data connectors, which bring logs into the workspace, failed 10 times across six connectors; the three analytics rules failed 7 of their 93 runs; automation rules failed 20 of 164. Only the one playbook ran cleanly. None of those failures raised an alert in the queue, because health isn't an alert: it's a record, and someone has to read it. Twenty failed automation runs, in particular, are twenty times the SOC's own response logic didn't do what someone designed it to do, and the incidents those runs should have touched carry no mark of the failure.
The connector failures are worth one more query, because a failing connector means missing data, and missing data means detections that can't fire.
SentinelHealth
| where SentinelResourceType == "Data connector" and Status != "Success"
| summarize Failures = count() by SentinelResourceName
Every connector failure came from one source: the Palo Alto Networks firewall at the edge of Northgate's network. For as long as a failure lasts, rules that read firewall logs work on a gap, and nothing in the queue says so. Ten failures in a month is a pattern, not an accident, and a pattern in a health record is exactly the kind of thing improvement exists to notice and send to whoever owns the connector. Module 2 teaches how to write rules that notice their own data going quiet; Module 11 teaches how to measure it across the whole SOC. The firewall is a good example of why it matters: firewall logs are where an investigator looks for a compromised device calling out to an attacker's server, and a gap in them is a gap in exactly the evidence that would show it.
Three readings of detection are common among people new to the work, and the month contradicts each.
The first row is the one to keep in mind for the whole course, because it's the one that fails silently. An alert is evidence that something the SOC decided to watch for happened; the absence of an alert is evidence of nothing, unless the SOC knows its detection covered the activity and its data was arriving. Most of the attacks in Northgate's month were caught by some alert somewhere; the parts that weren't were exactly where detection had an edge, a missing table, a rule that didn't exist, or a signal nobody had decided to watch.
Triage
Is it real?Triage is the first question asked of every incident: is this real, and how urgent is it? It's fast by design, because the queue doesn't wait, and it ends in a verdict. When an incident is closed, the verdict is recorded as a classification, and counting the month's closures by classification shows what triage decided.
SecurityIncident
| summarize arg_max(TimeGenerated, Status, Classification) by IncidentNumber
| where Status == "Closed"
| summarize Incidents = count() by Classification
Of 437 closures, 308 carry no classification, and the other 129 split into 31 true positives, 31 benign positives and 67 false positives. The blank row needs reading before anything else, because it's two different things. Two hundred and eighty of the blanks are incidents Microsoft XDR resolved on its own, low-severity incidents and identity risks users cleared themselves, which close without a verdict by design; the other 28 were closed by the SOC, by analysts or a tuning rule, and left blank. Among the verdicts that were recorded, false positives were the largest group by far, more than twice the true positives: most of what the SOC judged was an alert that shouldn't have fired. The 31 true positives are the month's real attacks, counted incident by incident, and several of them belong to the same intrusions from Section 0.1. The three verdicts mean different things, and the difference matters more than new analysts expect.
What a classification says
Three verdictsThe benign positive is the verdict people most often get wrong, usually because it sounds like an excuse rather than a finding. It says the alert did its job: the activity it describes really happened, and it was allowed, such as an administrator using a tool the rule watches for. A false positive says the rule was wrong. The distinction also protects the people involved: calling an administrator's legitimate work a false positive hides the fact that the rule caught exactly what it was built for, and calling a broken rule's alert a benign positive hides the fact that the rule needs work. The two call for opposite fixes: a benign positive might need an exception for that administrator, while a false positive needs the rule itself changed. A SOC that records them interchangeably can't tell its engineers which rules to fix, and the difference is plain set side by side.
Benign and false recorded
interchangeably
The engineer can't tell
which rules to fixBenign positive: add an
exception for the admin
False positive: change
the rule
Each verdict points at a fixThe right-hand pane is why classification is a detection function's input as much as triage's output.
The blank row is the costliest where a person or the SOC's own rule left it. Twenty-eight incidents the SOC closed carry no verdict, so nothing can be learned from them: not whether the rule was right, not whether the activity was real. Module 1 teaches why a classification is required on closure and what each value commits the analyst to.
Investigate
What happened, across tablesInvestigation is what happens after triage says an incident might be real. It asks what happened, how far it reached, and what the attacker can still do, and it answers by reading records across many tables, not only the alerts. The domain compromise from Section 0.1 is a good example, because its eight alerts describe four different stages of an attack.
let Ids = SecurityIncident
| where IncidentNumber == 70453
| summarize arg_max(TimeGenerated, AlertIds)
| mv-expand Id = parse_json(AlertIds)
| project Id = tostring(Id);
SecurityAlert
| where SystemAlertId in (Ids)
| summarize Alerts = count() by Tactics
Five of the eight alerts are about credential access, stealing passwords and tickets; the others are about moving between systems, persisting, and raising privileges. Those tactic names come from MITRE ATT&CK, the public catalog of attacker techniques that Defender tags its alerts with, and they're one of the first things an investigator reads, because they say which stage of an attack each alert describes.
What the alerts don't say is what happened between them, or after them. Section 0.1 showed that this incident left out the attacker's cloud sign-in the next morning, which only the sign-in logs record. That's the shape of most investigations: the alerts point at the start, and the answer is in tables no alert mentions. Module 7's playbooks teach how to move from an incident's alerts to the records around them in a fixed order, so nothing is left to memory. The next place to look depends on what raised the alert.
Where an investigation looks next
Beyond the alertsThe last row is the one that found the domain compromise's cloud sign-in, and it's the one most often skipped under time pressure, because the incident page makes the alerts feel like the whole story.
Investigation is also where the analyst's written record is made. A finding nobody wrote down can't be handed to the next shift, reviewed by a lead or used by an engineer, which is why Module 8 treats documentation as part of the investigation rather than a report written afterwards.
Respond
Contain, recover, recordResponse is what the SOC does about a real attack: contain it so it can't spread, recover what it damaged, and record what was done. In a Microsoft tenant some response is automatic. Defender's automatic attack disruption can contain a user or a device on its own when it's confident an attack is under way, and it records each action.
DisruptionAndResponseEvents
| summarize Events = count() by ActionType
Seven containment actions in the month, each blocking a contained user from logging on, opening files over the network or making remote calls, the three routes an attacker with a stolen account would use to move between systems. Nobody in the SOC clicked anything for these; the platform acted, and the records say what it did. The actions are recorded against the contained account and the time, so reading them takes one query. An analyst who didn't know to look would never learn that the attacker's access had already been cut, and might spend an hour containing something already contained.
Response by the SOC's own automation also shows in the health records from Section 02: 164 automation rule runs, 20 of them failed, and 62 playbook runs. Those are responses too, and the failures are responses that didn't happen. Set side by side, the picture of response most people arrive with and the one the records show look quite different.
Response is what analysts do
after an investigation
Slow, careful, manualResponse is also what the
platform did on its own:
7 containment actions,
automation and a playbook
Much of it before anyone lookedThe right-hand pane doesn't make the analyst's response less important; it changes where it starts. An analyst responding to an incident first reads what the platform and the automation already did, then decides what's left. Module 10 teaches both the automation and the platform's disruption, and how to check each one's work. Automatic response is powerful and fallible in the same way the agent is: it acts on what it can see, and a containment that blocked the wrong account, or missed the right one, is found only by someone reading what it did.
Improve
Health, cost and intelligenceImprovement is the function that turns a month of work into a better next month. It reads the SOC's own records, health, cost, verdicts, intelligence, and changes detection, triage and response because of what they show. Its inputs exist whether or not anyone reads them; the cost of the data the SOC collects is one of them.
Usage
| summarize GB = round(sum(Quantity) / 1024, 1) by DataType
| top 5 by GB desc
The five largest tables in the workspace held more than 130 GB between them in the month, led by the non-interactive sign-in logs, the records of applications and tokens signing in on users' behalf. Every gigabyte costs money, and the question improvement asks is whether each table earns its cost by feeding detections or investigations. The non-interactive sign-in table is a good test: it's the largest, most of it describes ordinary token refreshes, and it's also where Module 13 found nine uses of a stolen token that the interactive sign-in log never showed. A table that looks like noise in a cost report can be the one an investigation needs most, which is why cost decisions belong to people who know what each table is for. Module 11 answers it table by table.
Three readings of improvement keep it from happening, and each is common in busy SOCs.
The other improvement records have appeared already in this sub and the last. The connector failures say where detection has gaps. The SOC's 28 blank classifications say where triage leaves no lesson. The agent's reversed closures and the tuning rule's hidden reports, from Section 0.1, say where automation needs checking. Each is a record the SOC already has, and each points at a specific change someone could make this month. Module 12 adds threat intelligence, which the SOC can use to improve detection before an attack rather than after it, and Module 9 adds hardening, which removes some attacks from the queue entirely by making them impossible.
How They Depend on Each Other
A loop, not a lineThe four functions aren't independent teams working in sequence. Each one's output is the next one's input, and improvement feeds the first, which makes the whole thing a loop.
The loop explains why a weakness anywhere shows up everywhere. Noisy detection overloads triage; careless triage leaves investigation with bad leads and improvement with blank classifications; unchecked response leaves attackers where the records say they were removed; and without improvement, next month repeats this one. A SOC that fixes one function in isolation usually finds the problem has moved. Each weakness shows up somewhere other than where it starts.
A weakness, and where it shows up
Knock-on effectsThe right-hand column is what a SOC lead usually sees first, and the left-hand column is where the fix has to go. Hiring more analysts for the first row treats the symptom; tuning the rules that raise the false positives treats the cause, and it's cheaper. The same pattern runs through the others, which is why the course keeps asking where a problem started rather than where it was noticed.
It also explains why the course teaches the functions in the order it does, and why the order isn't a ranking of importance. Detection comes first because it sets the load for everything after it; triage and investigation next because they produce the verdicts the rest depends on; response after them; and improvement last, because it needs all three to have left records worth reading.
Where Each Function Is Taught
Mapping the courseThe course gives each function its own modules, and several modules serve more than one function. Mapped out, the course looks like this.
Where each function is taught
ModulesRead the map as a reference, not a sequence, since each module teaches its subject completely and can be read on its own, and come back to it whenever a module seems to wander from its title, because the wandering is usually the module serving a second function: most modules can be read on their own, and the functions they serve overlap. What the map does show is that improvement gets as many modules as detection. That's deliberate, and it reflects what separates SOCs that get better from those that stay busy: the habit of reading their own records and acting on them. A SOC that only detects, triages and responds does the same work next month, with the same gaps; one that also improves does slightly different work, and finds slightly more. Over a year, the difference between the two compounds: fewer false positives, fewer blank classifications, fewer silent gaps, and analysts whose time goes on the incidents that matter.
Practice
Find each function's records in the lab.
Four functions, four queries
In this course's lab; read-only.
- Read SentinelHealth for failures, and name the connector that failed.
- Count closures by classification, and explain what the blank row costs.
- List the containment actions the platform took on its own.
- Find the five largest tables by volume, and guess which detections read each.
Keep your guesses for the last step; Module 11 shows how to check them.
The next sub, Section 0.3, looks at where SOCs commonly fail, and finds most of the failures in Northgate's month were failures of the fourth function.