Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
0.5 The Detection-and-Response Pipeline
Introduction
Every detection a SOC makes passes through the same stages. Something happens in the estate; a product or a rule notices and raises an alert; the alert becomes an incident; someone is given the incident; someone works out what happened; and the incident is closed with a verdict, and responded to if it was real. That sequence is the detection-and-response pipeline, and it's the backbone of this course: each module teaches one or more of its stages. Nothing in this sub assumes you've worked in a SOC; each stage is explained where it first appears, with the table it's recorded in. It follows one real incident from Northgate's month through all six stages, reads the timestamp each stage left, and then measures the same stages across every true positive in the month, to show where the pipeline is fast, where it's slow, and where incidents fall out of it altogether.
The numbers above the arrows are the point. Each stage writes a timestamp, so each gap can be measured, and the gaps are very different sizes: seconds between alert, incident and owner; an hour or two between activity and alert, and between owner and work; most of a day between work and closure. Knowing which gap is which is what lets a SOC improve the right stage instead of the most visible one. The most visible gap is usually the last, time to close, because it's the one reports lead with; it's also the one that least needs shortening.
Six Stages
From activity to closureThe six stages are worth naming precisely, and in order, because each has its own record and its own failure mode. In Microsoft's tools they map onto the tables this course reads throughout, from its first module to its last.
The six stages
From activity to closureThe first stage happens in the estate, and its record is whatever table the activity lands in: sign-ins, device events, mail. It's the only stage the SOC doesn't control, and the only one that happens whether or not anyone is watching. The second is the SecurityAlert table, with a start time for the activity and a time for the alert itself. The third, fourth and fifth are all in SecurityIncident, which writes a new row every time the incident changes. The sixth is a classification on the final row, and, if the incident was real, response actions recorded in their own tables.
Written as a list of tables, the stages look like this, and every table in the list is one this course queries, most of them in more than one module.
Where each stage is recorded
TablesKnowing the table for each stage is what makes the stage measurable: every gap in this sub is the difference between two timestamps in two of those tables. Each stage can fail on its own. Activity can happen where no product is watching. An alert can be raised and grouped into an incident that nobody is given. An owner can be given an incident and never open it. An investigation can stop at the incident's alerts and miss what happened outside them. A closure can record no verdict, or a wrong one. Sections 0.3 and 0.4 met most of these as failures; this sub meets them as stages, in order.
One Incident
Incident 83098, change by changeIncident 83098 is ordinary, which makes it a good one to follow: a password spray against one account on 23 February, raised by Defender for Identity and closed by an analyst as a true positive. Its history is four rows in the incident table.
SecurityIncident
| where IncidentNumber == 83098
| project TimeGenerated, Status, Owner = tostring(parse_json(Owner).assignedTo),
Classification, ModifiedBy
| sort by TimeGenerated asc
Four rows, four changes, each with the name of whoever or whatever made it, and each with a timestamp to the second. Sentinel created the incident at 12:34:17. Four seconds later, an automation rule tagged it as an identity incident and assigned it to Priya Sharma. At 13:53 she set it Active, which is the first sign of a person working on it. At 05:37 the next morning she closed it as a true positive.
Nothing in the rows is unusual, and that's what makes them worth reading closely: most incidents in any SOC have exactly this shape, created by the platform, assigned by a rule, worked and closed by a person. Read as a timeline, the same four rows show which stage each change belongs to.
The two marked lines are the ones people confuse. The automation rule's change at 12:34:21 gave the incident an owner; the analyst's change at 13:53 started the work. On a dashboard both might count as the incident being picked up, and the gap between them, an hour and 19 minutes, disappears. That gap is the stage where Section 0.4 found most of Northgate's queue stuck: owned, and never started. For incident 83098 it closed within the afternoon, because Priya Sharma reached it; for 43 others that month it never closed at all. The rows look identical up to that point, which is why the owner field alone says so little about whether an incident is being worked.
From Activity to Alert
When did it actually begin?The incident was created at 12:34, but the activity it describes started earlier. The alert records both times: when the activity began and ended, and when the alert was raised.
let Ids = SecurityIncident
| where IncidentNumber == 83098
| summarize arg_max(TimeGenerated, AlertIds)
| mv-expand Id = parse_json(AlertIds)
| project Id = tostring(Id);
SecurityAlert
| where SystemAlertId in (Ids)
| project StartTime, EndTime, TimeGenerated, AlertName, ProviderName,
CompromisedEntity
The alert's own record, in SecurityAlert, carries the activity's time as well as its own, which is what makes this stage measurable at all. The same query works for any incident: take its alert IDs, and read the alert rows they point to. The spray began at 10:45:39 and ran until 12:29:49; Defender for Identity raised the alert at 12:33:39, four minutes after the activity ended and an hour and 48 minutes after it began. Whether that's fast or slow depends entirely on which of the two times you measure from. Most reports measure from the alert, because that's the first time the SOC could have known; the account's owner, and the attacker, measured from the start.
Alert raised at 12:33
detection took 4 minutes
FastActivity began at 10:45
alert raised at 12:33
the spray ran for nearly two
hours before anyone could know
As fast as the pattern allowsBoth panes are true, and they answer different questions. Measured from the end, detection was quick: the product decided within minutes of the pattern completing. Measured from the start, the account was under attack for nearly two hours before anyone could know, and that's the number that matters to the person whose account it was. A password spray is a pattern, and a pattern takes time to become visible; some detections can't fire earlier without firing wrongly, which is a trade Module 2 teaches how to make deliberately.
Every alert carries three times, and the difference between them is what this section has been reading.
Three times on every alert
SecurityAlertThe general lesson is to read StartTime as well as TimeGenerated. An alert's own time is when a product decided; the activity's time is when the attacker acted, and an investigation's timeline has to start there. The EndTime matters too: an alert whose activity ended hours before it was raised describes something finished, and one whose activity is still running describes something the SOC can still interrupt.
From Alert to Owner
Seconds, and what they hideThe next two stages took seconds. The incident was created 38 seconds after the alert, and the automation rule assigned it four seconds after that. This is the part of the pipeline software does, and it does it quickly and consistently.
The rule that did the assigning has a health record of its own, and it's worth reading once, because a stage that runs in seconds usually runs unwatched.
SentinelHealth
| where SentinelResourceName has "Tag and assign"
| summarize Runs = count(), Failures = countif(Status != "Success"),
First = min(TimeGenerated), Last = max(TimeGenerated)
Eighty-two runs in the month and no failures: this stage is working, and it's the one stage in this sub with nothing to fix. Its only weakness is what it covers: it assigns identity incidents, and the queue's untouched incidents are mostly kinds it doesn't match. Speed here is almost free, and it can mislead. An incident assigned in four seconds looks handled on any report that counts assignment. What the assignment doesn't say is whether the owner was available, whether they had fifteen other incidents, or whether they would open this one today. Section 0.4 found 43 incidents owned and never started; each of those was assigned just as quickly as 83098, and then waited.
Three readings of the pipeline's timestamps are common, and incident 83098 shows why each is wrong.
The middle row is the one that matters most for this stage. Assignment is a promise that someone will look, and the pipeline only measures the promise unless the SOC also measures the first person's change. Module 1 teaches how the queue turns assignment into work, and Module 10 teaches which assignments automation should make.
From Owner to Answer
The stage that takes the timeAn hour and 19 minutes after assignment, Priya Sharma set the incident Active, and closed it as a true positive at 05:37 the next morning, fifteen hours and 44 minutes later. The incident's own rows don't say what she did in that time, which is itself a finding: Module 8 found that none of the month's incidents carries a comment, so the investigation left no record of its reasoning.
The close time looks long, and it probably isn't slow work, as the sign-ins show. Investigation is the stage where an analyst reads the records around an alert: the account's sign-ins, its mailbox, its devices, what else the attacker's address touched. A password spray against one account needs the account's own activity checked for the hours around it, and that can't be rushed without missing something. For e.bevan, the first thing the investigator reads is the account's sign-ins that day.
SigninLogs
| where UserPrincipalName == "e.bevan@ne.com"
| where TimeGenerated between (datetime("2026-02-23T10:00:00Z")
.. datetime("2026-02-24T06:00:00Z"))
| project TimeGenerated, IPAddress, Location, AppDisplayName, ResultType
| sort by TimeGenerated asc
All three succeeded, and none was scored risky. Three successful sign-ins, two of them from an address in Sweden, one during the spray itself and one after, and one from the UK. That's exactly the kind of evidence the investigation has to weigh: travel, a VPN, or an attacker who got in. The records alone don't say which; the analyst's job is to find out, and that takes time the pipeline should protect.
Most of that elapsed time was also overnight: an incident closed at 05:37 was closed at the start of a shift, not after fifteen hours of continuous work. Set side by side, a closure made quickly to meet a target and the one that actually happened show what the time bought.
Closed in 4 hours
sign-ins not read
verdict: true positive
FastClosed in 16 hours
sign-ins from Sweden weighed
verdict: true positive,
with the reasoning
SlowerThat's why the closure stage is the wrong place to look for speed, and why a review of it should read the reasoning rather than the clock. It's the stage where the SOC's judgment happens, and squeezing it produces faster wrong answers. The stages worth speeding up are the ones before it, where nothing is being judged, only waited for.
The Month's Pipeline
Every true positive, stage by stageOne incident shows the stages; the month's true positives show whether its timings were typical. True positives are the right population for this, because they're the incidents where every stage mattered. The query below takes each true positive the SOC closed, joins it to its first alert, and computes the median of each gap.
let Life = SecurityIncident
| where not(ModifiedBy == "Microsoft XDR")
| summarize Created = min(todatetime(CreatedTime)),
Active = minif(TimeGenerated, Status == "Active"),
Closed = minif(TimeGenerated, Status == "Closed"),
arg_max(TimeGenerated, Classification, AlertIds) by IncidentNumber
| where Classification == "TruePositive"
| extend Id = tostring(parse_json(AlertIds)[0]);
SecurityAlert
| join kind=inner Life on $left.SystemAlertId == $right.Id
| extend DetectMin = toint((TimeGenerated - StartTime) / 1m),
ToIncidentSec = toint((Created - TimeGenerated) / 1s),
ToActiveMin = toint((Active - Created) / 1m),
ToCloseH = toint((Closed - Created) / 1m) / 60.0
| summarize Incidents = count(), DetectMin = percentile(DetectMin, 50),
ToIncidentSec = percentile(ToIncidentSec, 50),
ToActiveMin = percentile(ToActiveMin, 50),
ToCloseH = round(percentile(ToCloseH, 50), 1)
Thirty-one true positives, and incident 83098 was close to typical. The median alert came four minutes after its activity's start time, the median incident thirty seconds after its alert, the median first work 38 minutes after the incident, and the median closure eleven and a half hours after creation. The detection median is shorter than 83098's because most true positives are single events rather than patterns: a sign-in or a message is visible the moment it happens. The median to first work, 38 minutes, is shorter too, because many true positives were reported phish the agent had already flagged as real, which an analyst picks up quickly.
Set against those medians, incident 83098 was slower than typical at every stage a person or a pattern controls, and typical where software worked, which is the usual way one incident differs from the median.
Incident 83098 against the month
Its gaps and the mediansNone of its gaps is alarming on its own. The shape is the same as the single incident's: seconds where software works, minutes to hours where people start, hours where people investigate. It's a healthy shape for the incidents that make it through: software does the fast stages, people spend their time where judgment is needed, and the slowest stage is the one that should be slow. A SOC with this shape doesn't need to speed up its analysts; it needs to make sure every incident reaches one. The next section asks about the ones that don't.
Where the Pipeline Leaks
Incidents that never reach a personMedians, rather than means, are the right summary for skewed times like these, because one incident closed after a week would drag a mean far from anything typical, as Module 11 shows. The medians above are computed over incidents that were closed, so they can only describe incidents that made it through every stage. A pipeline can be fast for those and still lose many others along the way, and counting how many reach each stage shows where.
SecurityIncident
| summarize arg_max(TimeGenerated, *) by IncidentNumber
| where not(Status == "Closed" and ModifiedBy == "Microsoft XDR")
| summarize Incidents = count(), Touched = countif(isnotempty(FirstModifiedTime)),
Closed = countif(Status == "Closed"),
WithVerdict = countif(Status == "Closed" and Classification != "")
Of 260 incidents the SOC dealt with, 201 were touched, 157 closed and 129 closed with a verdict. The pipeline leaks at every stage after assignment, and the largest leak is the first: 59 incidents never touched by anyone, nineteen of them High. Those 59 have no time to first work, no time to close and no verdict, so no median can include them, and a pipeline report built from medians alone would never mention them.
The counts also say, stage by stage, where the other, smaller leaks are: 44 touched incidents weren't closed, and 28 closures carry no verdict. Each is a stage the pipeline started and didn't finish. This is the most common way SOCs misjudge their own pipeline: measuring the speed of what gets through and not the volume of what doesn't. The proposal below is the usual first reaction to a long closure time; make the call before reading where it resolves.
Judgment Call
Improvement decisionProposal, from Northgate's SOC lead
"Our median true positive takes 11 hours to close. The fastest win is closing faster: set a target of four hours and hold analysts to it."
Across the month's 31 true positives: detection a median 4 minutes after the activity began, the incident 30 seconds after the alert, first work 38 minutes after that, closure a median 11.4 hours after creation. Separately, 59 of the 260 incidents the SOC dealt with were never touched at all.
What is your call?
The call rests on the counts, not the medians. A closure target makes the fast part of the SOC faster at the expense of its judgment, and leaves the 59 exactly where they were. Getting every incident touched, Highs first, is a smaller change in habits and a much larger change in what the SOC actually sees. It's also measurable with the same query as the leak itself: next month, the Touched column should be closer to the Incidents column, and the difference should hold no Highs.
After Closure
Response, and the loop backClosure records a verdict; it doesn't remove an attacker. For a true positive, the pipeline continues into response: containment, recovery, and the changes that stop the same attack working twice. Some of that is done by people, and some by the platform itself, which records its own actions.
DisruptionAndResponseEvents
| summarize Events = count(), First = min(TimeGenerated), Last = max(TimeGenerated)
by ActionType
Seven automatic containment actions, all in the early hours of 12 March, blocking a compromised user from logging on, opening files over the network and making remote calls. They happened within eight minutes of each other, at 02:34 in the morning, with no analyst involved, and they belong to the same stage as an analyst's password reset: the response that follows a real detection. Module 10 teaches how the platform decides to act on its own and how a SOC checks what it did.
For incident 83098, the records show a classification and nothing else, which is common: no comment, no containment the platform recorded, no note on the detection. What a closed true positive should lead to is worth writing down, because each item is a record someone can later check.
What a closed true positive should lead to
After closureThe last two rows are where the pipeline starts to loop, and they're the two most often skipped, because by the time an incident is closed the next one is waiting. The last part of the pipeline loops back to the first. A true positive is evidence that a detection worked; a false positive is evidence that one needs tuning; a missed attack is evidence that a detection is missing. Recording which is which, and acting on it, is the improvement function from Section 0.2, and it's what turns the pipeline from a line into a loop. Mapped to the course, the stages look like this.
The fourth rung is the one that closes the loop, and the one most SOCs leave for later because nothing in the queue demands it, which is why the course's last modules are about hardening, measurement, threat intelligence and AI: each takes what the earlier stages recorded and uses it to change the first stage. Every module you read after this one sits somewhere on that ladder, and knowing where it sits helps you read it. When a detection module feels far from an investigation module, it's because they teach different stages of the same pipeline; when an incident in one module turns up again in another, it's the same incident at a later stage.
Practice
Follow one incident of your own choosing through the pipeline.
One incident, six timestamps
In this course's lab, or read-only in your own workspace.
- Pick a closed true positive and list its incident history, change by change.
- Find its alert, and compare StartTime with TimeGenerated.
- Separate the assignment from the first person's change.
- Run the month's medians, then count the incidents they leave out.
- Look for any response the platform took on its own.
Write down which gap in your incident was longest, and whether anything was being judged during it, or only waited for.
The next sub, Section 0.6, sets out what this course builds, module by module, and what you'll have at the end of it.