Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
0.3 Where SOCs Fail
Introduction
SOCs rarely fail in the way people imagine, with an alarm ignored or a screen nobody was watching. They fail quietly, in ways that look like normal operation from the inside: a queue that's always a little too long, verdicts left blank because the next incident was waiting, a data feed that stopped for a few hours, a rule or an agent whose decisions nobody went back to check. None of these raises an alert, and all of them leave a record. This sub finds six common failures in Northgate's month, shows the record each one left, and names the module in this course that addresses it. The point isn't that Northgate's SOC is bad; its month is ordinary, and ordinary is exactly where these failures live. Each section below is short on blame and long on records, because the records are what make a failure fixable.
The diagram's right-hand box is the sub's conclusion stated first. Each failure on the left has its own symptoms, and each was visible in records the SOC already held. What they share is that nobody was reading those records, or nobody had been given the job of reading them, which is a failure of the fourth function from Section 0.2, improvement, showing up in the other three. That's good news for anyone starting in a SOC: the failures are findable by anyone willing to read, and finding them needs queries rather than seniority.
Failing Quietly
Six failures in an ordinary monthBefore looking at each failure, it helps to see them together, because together they show a pattern no single failure does, with the piece of evidence that shows each one. Every row below comes from a query in this sub or in Sections 0.1 and 0.2, and every one can be rerun in the lab.
The month's failures, and where each shows
The evidenceNone of these six would make a SOC lead's weekly report unless someone went looking, and that's the defining feature of how SOCs fail. A weekly report usually counts what was done: incidents closed, alerts handled, time to respond. Each of the six failures is something not done, and a report of what was done can't show what wasn't. A breach is loud; the conditions that let a breach go unnoticed are quiet. A queue with a tail, a connector that fails at night, a verdict left blank: each is a small thing on the day, and each compounds across a month into a SOC that's busy, feels productive, and misses things.
The six aren't independent, either, and the first three feed each other in a loop.
The marked lines are where the loop becomes self-sustaining, and where breaking it is cheapest. Once blank verdicts stop the tuning, the false positives never fall, and the queue stays long however hard the analysts work. Most failing SOCs have several of these at once, reinforcing each other, which is why fixing one in isolation rarely helps for long.
The rest of this sub takes them one at a time, in the order an incident meets them: arriving, waiting, closing, and the data, automation and scope around all three. For each, it shows the record, explains why the failure is easy to miss, and points to where the course teaches the fix.
Too Much to Read
Where the noise comes fromEvery SOC complains about noise: alerts that turn out to be nothing, taking time that real incidents need. Section 0.2 found 67 false positives among the 129 closures that carry a verdict. The useful question isn't how many, but where they came from, and grouping the false positives by the alert that raised them answers it.
let FalsePositives = SecurityIncident
| summarize arg_max(TimeGenerated, Status, Classification, AlertIds)
by IncidentNumber
| where Status == "Closed" and Classification == "FalsePositive"
| mv-expand Id = parse_json(AlertIds)
| project IncidentNumber, Id = tostring(Id);
SecurityAlert
| join kind=inner FalsePositives on $left.SystemAlertId == $right.Id
| summarize FalsePositives = dcount(IncidentNumber) by AlertName
| top 5 by FalsePositives desc
The answer is lopsided. Fifty-two of the 67, more than three quarters, were the same kind of incident: a user reporting a message as phishing, and the message turning out to be harmless. The other fifteen are spread thinly across Microsoft's identity and cloud alerts, none contributing more than four. The noise at Northgate isn't a noisy rule; it's careful users reporting ordinary mail, which is a very different problem with a very different fix.
That's the most common mistake SOCs make about noise: assuming it comes from their detections and tuning the wrong thing. Here, tuning analytics rules would change almost nothing, while the Phishing Triage Agent already closes most of those reports in minutes and a short notice to staff before routine internal mail would prevent some of them. Three readings of noise are common, and the month contradicts each.
The second row is the one that matters most. Noise isn't only wasted time; it's time taken from somewhere, and in a queue with a tail, the somewhere is the incidents nobody reaches. Module 11 measures noise by source so the SOC fixes the right thing, and Module 2 teaches how to tune a rule without silencing what it catches. The third row matters because the quickest way to cut noise, turning off the loudest source, is also the quickest way to stop seeing that source's real attacks; the reported-phish alerts that make up most of the noise also carried three real phishing campaigns in the month.
What Nobody Picks Up
Fast on some, silent on othersThe queue's tail is the failure that hides best, because the usual measure of a SOC's speed can't see it. That measure is time to first touch: how long an incident waits before a person first opens it. Computed over every incident a person touched, Northgate looks fast.
SecurityIncident
| summarize Created = min(CreatedTime),
FirstHuman = minif(TimeGenerated, ModifiedBy endswith "@ne.com") by IncidentNumber
| where isnotnull(FirstHuman)
| extend Hours = toint((FirstHuman - todatetime(Created)) / 1m) / 60.0
| summarize Incidents = count(), MedianHours = round(percentile(Hours, 50), 1),
P90Hours = round(percentile(Hours, 90), 1)
A person first touched half of those 113 incidents within 36 minutes, and nine in ten within two hours. For a SOC of Northgate's size, that's quick. The query takes the earliest change made by a person, identified by a Northgate address, so the agent's and the tuning rule's changes don't count as a first touch; a first touch by software is a different thing, and Module 11 measures it separately. That's a good figure by most standards, and a report built on it would describe a responsive SOC. It's also computed only over incidents someone touched, which by definition excludes the 102 nobody ever did. Put side by side, the two descriptions of the same month read very differently.
Median first touch: 36 minutes
90% touched within 2 hours
A fast SOC113 incidents touched quickly
102 never touched at all
A fast SOC with a tail
nobody measuresThe right-hand pane is the honest one, and it isn't a criticism of the analysts, who were quick on everything they reached. It's a criticism of the measure, which can only report on work that started, and of any report that uses it alone. A SOC can be fast on everything it picks up and still leave almost all of its open queue untouched, and the speed figure will never show it, because an incident with no first touch has no time to measure. The oldest of Northgate's untouched incidents had waited since the first day of the month. What the tail holds is worth seeing, because it isn't random.
let Waiting = SecurityIncident
| summarize arg_max(TimeGenerated, Status, CreatedTime, AlertIds)
by IncidentNumber
| where Status == "New"
| mv-expand Id = parse_json(AlertIds)
| project IncidentNumber, CreatedTime, Id = tostring(Id);
SecurityAlert
| join kind=inner Waiting on $left.SystemAlertId == $right.Id
| summarize Waiting = dcount(IncidentNumber), Oldest = min(todatetime(CreatedTime))
by ProviderName
| sort by Waiting desc
Sixty-three of the 102 come from cloud-application alerts and 33 from identity alerts. Three come from Defender for Endpoint, dated from 10 March, the week ransomware reached a laptop, and nobody opened them. That pattern says the tail isn't bad luck; it's a category of work the SOC has, in effect, stopped reading, without anyone deciding to.
The fix isn't only more speed. It's measuring the tail directly, as Module 11 does, and deciding which kinds of incident the SOC will leave untouched, rather than leaving that to chance. Module 1 teaches the triage that makes those decisions quickly, and Module 13 asks whether an agent could read part of the tail.
Verdicts That Teach Nothing
Twenty-eight blank closuresA classification is the verdict an incident closes with, and it's what connects the queue to improvement: a false positive tells an engineer which rule to fix, a true positive counts toward the month's real attacks, and a benign positive records that the alert was right about something allowed. Section 0.2 found 308 closures with no classification. Who closed them is the first thing to know.
SecurityIncident
| summarize arg_max(TimeGenerated, Status, Classification, ModifiedBy)
by IncidentNumber
| where Status == "Closed" and Classification == ""
| summarize Incidents = count() by ClosedBy = case(
ModifiedBy endswith "@ne.com", "an analyst",
ModifiedBy startswith "Alert tuning", "the tuning rule", ModifiedBy)
Two hundred and eighty were the platform's automatic resolutions, which close without a verdict by design and aren't a failure. The SOC's own blanks are the other 28: twenty closed by analysts, and eight by the tuning rule that auto-resolved reported phish in the first days of the month. Twenty out of the analysts' 98 closures is about one in five. The two groups fail for different reasons. Neither group is careless. The analysts' blanks are usually time pressure: the incident was clearly done, the next one was waiting, and the verdict felt like a formality. The tuning rule's blanks are by design: it closes without judging, which is exactly why it closed two real phishing reports without anyone noticing.
The analysts' twenty are the ones a SOC can change by habit, and the tuning rule's eight the ones it changes by configuration. Either way, the cost is the same, and it falls on three different people.
What a blank closure costs
Three lossesThe first row costs the engineer, the second costs whoever reports the SOC's performance, and the third costs the next analyst, who meets the same alert with no idea what the last person concluded. None of them is present when the blank is left, which is why the habit persists. A classification takes seconds to record and months to reconstruct, and a SOC that measures blank closures as a percentage of the month usually finds the number falls as soon as someone starts reading it. Module 1 makes classification a required part of closing, and Module 8 teaches what each verdict commits an analyst to writing down.
Blind Spots
A connector failing one run in threeA detection can only fire on data that arrives. When a data connector fails, rules that read its table run on nothing, and they do it silently: no alert says that an alert couldn't be raised. Section 0.2 found the Palo Alto firewall's connector behind every connector failure in the month. Its own record shows how bad it was.
SentinelHealth
| where SentinelResourceName == "PaloAltoNetworks"
| summarize Runs = count(), Failures = countif(Status != "Success"),
FirstFailure = minif(TimeGenerated, Status != "Success")
Ten failures in 31 runs, one run in three, starting on 6 March and continuing to the end of the month. Nobody fixed it in that time, which is the clearest sign that nobody knew, since a firewall feed is not something any SOC chooses to lose quietly. For more than a week, the firewall's picture of traffic leaving Northgate had holes in it, and any rule or investigation that relied on it was working from an incomplete record without knowing so. That includes the question an investigator asks first about a compromised device: did it call out to an attacker's server? The ransomware that reached a laptop on 12 March did exactly that, and any firewall record of that traffic falls inside the window when the connector was failing a third of the time.
The failure isn't that the connector broke; connectors break. It's that nobody heard. Health records are written whether or not anyone reads them, and the question every SOC should be able to answer is the one below.
reader a named owner for each connector and rule
check a scheduled query over SentinelHealth
route an incident, or a message to the ownerHealth nobody reads is a gap nobody knows about.The gate's foot is the whole lesson, and it applies to rules as well as connectors: the three analytics rules failed seven of their runs in the month, each failure a window in which the rule saw nothing. Northgate had the record, and nobody's job was to read it. Module 2 teaches rules that check their own data is arriving, and Module 11 makes health a measured part of the operation, with an owner for every connector. Owners are what turn a record into an action, and each kind of health record has a natural one.
Who reads which health record
OwnersNone of those owners needs to read every run. A daily query that lists failures by resource, sent to the owner of each failed row, is enough to turn a week-long gap into an afternoon's, and the query is short enough to write on the first day in a new SOC.
Unchecked Automation
Workers whose output nobody readsAutomation is supposed to reduce the analysts' load, and it does. It also introduces workers whose output needs the same scrutiny as a person's, and that scrutiny is easy to skip because automation doesn't ask for it. Northgate's month had three kinds of automated worker, and each had a problem nobody saw at the time.
Automation's unchecked month
Three workersThe tuning rule's two hidden phish, the twenty failed automation runs and the agent's three misses were all found later, by going back through records. None was found by the automation reporting its own problems, because automation doesn't report what it didn't do or got wrong. The agent's own figures, Module 13 showed, report its speed and volume and never its errors.
That's the failure: treating automation as finished work rather than as a worker, whose output is a claim like any analyst's. The fix is the same as for any worker, a check on a sample of its output and a record of what the checks find, and Modules 10 and 13 build exactly that for automation rules and for AI agents. It's worth building early, because automation's mistakes repeat at automation's speed. Three readings of automation keep SOCs from building it.
The last row is the one that surprises teams most, and the one Module 13 spends four sections on, because the agent's dashboard looks so complete: daily activity, time to triage, capacity used, all accurate. Complete about the wrong thing is still incomplete.
Incidents Drawn Too Small
The attack outside the incidentThe last failure is an investigation failure rather than an operational one, and it's by far the subtlest, because it happens to careful analysts working exactly as the incident page invites them to. An incident groups alerts, and an analyst investigating it naturally reads those alerts. The attack, though, doesn't know where the incident's boundary is. Northgate's domain compromise is the clearest example: eight alerts in one incident, and the most important evidence outside it.
Incident 70453
8 alerts, on-premises
one intrusion, one incident
Scoped by its alertsThe intrusion
8 alerts, on-premises
a cloud sign-in next morning
the same relay as February's
phishing
Scoped by the recordsTwo facts sit outside the incident on the right, and each came from a different table. The compromised administrator signed in to the Azure portal the next morning, which no alert in the incident records; and that sign-in came from the same relay address that took over a finance user's sessions two weeks earlier, in a phishing attack filed as separate incidents entirely. An analyst who scoped the investigation by the incident's alerts would have closed a domain compromise while the attacker was in the cloud, and treated one intrusion as two.
The address is the link, and one query shows it.
SigninLogs
| where IPAddress == "185.234.72.18"
| summarize SignIns = count(), First = min(TimeGenerated), Last = max(TimeGenerated)
by UserPrincipalName
The same relay served both attacks, two weeks apart, and nothing in either incident points to the other. An attacker reusing infrastructure is common, because infrastructure costs money and effort to replace, and it's one of the most reliable ways to link incidents a SOC has filed separately. The fix is a habit rather than a tool: after reading an incident's alerts, look for its accounts, devices and addresses everywhere else. Module 7's playbooks make that step routine, and Modules 3 to 6 teach the tables it reads.
One Cause
And where the course fixes each failureSix failures, and one pattern under all of them: every one was visible in a record the SOC already held, and none was noticed because nobody's job was to read that record. That gives a SOC lead a concrete place to start, which is usually not where the first instinct goes. The proposal below is the common instinct; make the call before reading where it resolves.
Judgment Call
Improvement decisionProposal, from Northgate's SOC lead
"We've found six failures in one month. Hire two more analysts: more people will read more of the queue and catch more of this."
The six failures: 67 false positives, 52 of them user-reported mail the agent handles; 102 incidents never touched while touched ones got a first look in a median of 36 minutes; 28 closures the SOC left unclassified; a firewall connector failing a third of its runs; unchecked automation; and an incident scoped to its alerts.
What is your call?
The judgment call is a real one in most SOCs, and there's no shame in the instinct at all, because more staff is the fix a budget holder understands. The resolution isn't against hiring; it's against hiring as the first fix for failures that aren't about staff. Most of the six persist with more people, because the people would read the same queue the same way. Giving each record an owner changes what gets read, and costs far less than a salary, and several of the failures, especially noise and the tail, shrink once their causes are addressed.
The same reasoning applies to buying tools, which is the third option's appeal and its weakness: a new tool is worth having when it closes a gap a record shows, and not before. Mapped to the course, each failure has a place where the fix is taught.
The ladder's foot is the reason this sub comes before the modules, and before any of the course's detections, playbooks or measures. Every fix starts with a record Northgate already has, which means every fix can start in a real SOC on the day someone decides to read it; none waits for a new product, a new budget or a new team. That also means none of the fixes is dramatic. They're small, repeated acts of reading: a health query every morning, a blank-classification count every week, a sample of automated closures every month. The course teaches each one where its subject is taught, so that by the end the habit is part of the work rather than an extra.
Practice
Find one of the six failures in the lab, and the record that shows it.
Six failures, six records
In this course's lab; read-only.
- Group false positives by alert, and name where the noise comes from.
- Compute time to first touch, then count the incidents it leaves out.
- Count blank closures by who made them.
- Find the connector with the most failures, and when they started.
- Pick the failure you would fix first, and name its owner.
Keep your choice and its owner for the last step; Section 0.4 shows how a SOC's maturity is judged by exactly these records.
The next sub, Section 0.4, turns the six failures into a way of measuring how mature a SOC's operation is.