Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
What Good Looks Like
Introduction
Endpoint operations gets measured, and most of the measures in common use can be improved without anything getting better. That is not cynicism about metrics; it is a property of which numbers are easy to collect.
There is one test that separates the useful from the decorative, and it is worth adopting before any of the individual measures below. If a number would change what you do on Monday, keep it. If it would only fill a slide, it is trivia, however true it is.
Apply it to the audience as well as to the number. The same measure can be a decision for one reader and trivia for another: alerts per incident should change what a detection engineer does this week and means nothing to a board, while retrospective detection time by incident type is exactly the number a board can act on and is too coarse to guide a tuning session. A pack that serves both audiences with one table serves neither well.
You will finish able to say what each common measure is actually counting, which of them can be improved by suppressing rather than by detecting, and how to state a detection time that describes the adversary rather than your own queue.
Scenario
A monthly pack shows six endpoint metrics and every one of them improved. Volume down 43 per cent, detection time down 71 per cent, false positives down 27 points, coverage up 16. The team is congratulated. Three tuning rules were deployed in July, the false positive rate is calculated only on alerts that reached an analyst, and the quarter's one real incident had three days of activity in the logs before anything fired.
Where You Start the Clock
The single most consequential choice in SOC measurementDetection time is the headline number in almost every security operations pack, and it is three different measurements wearing one name.
The distinction that matters
Activity to alert measures how fast your rules are. Alert to acknowledgment measures how fast your queue is. First evidence to confirmation measures how long the adversary had.
Most packs report the second and call it detection time. It improves when triage gets faster and does not move when detection gets worse, which is how a pack shows detection improving while detection degrades.
Drawn against a single incident, the three become obvious and the gap between them is the number nobody is reporting.
Figure EO0.8a. The orange bar improves when triage gets faster. It does not move when detection gets worse, which is why a pack can show detection improving while detection degrades.
Almost every pack reports one of the shorter bars because those can be computed automatically from ticket data. Nothing has to be looked up, no incident has to be reopened, and the number appears in an export. The measure that is easy to collect wins by default, which is the general reason SOC measurement drifts toward describing the SOC rather than the threat.
The version worth reporting is the long red bar, and it cannot be computed from alert timestamps. It requires going back into the logs after each confirmed incident, finding the earliest indicator, and measuring forward from there. That is sometimes called retrospective detection time, it takes twenty minutes per incident, and it is the only version that describes the adversary rather than your own responsiveness.
Twenty minutes per incident is a real cost and the reason most teams do not do it. Do it for the confirmed incidents only, which on a normal estate is a handful a quarter, and the number you get is worth more than the one you have been reporting.
There is a second reason to do it that has nothing to do with reporting. Working backwards from a confirmed incident to its earliest trace is the most reliable detection engineering input available: whatever that first indicator was, either a rule should have fired on it and did not, or nothing was watching that behavior at all. Every retrospective calculation therefore produces a candidate detection as a by-product, which makes the twenty minutes considerably easier to justify.
Record all three numbers rather than replacing one with another. Activity to alert tells you about your rules, alert to acknowledgment tells you about your queue, and the total tells you about the intrusion. They diagnose different problems and a team that reports only the total cannot tell which half to fix.
The Clock the Adversary Is Running
Which is why the units matterThe reason the starting point matters is that the other side moves fast, and the published figures have shifted sharply in the last few years.
Median ransomware dwell time has been measured at under twenty-four hours, with lockers deployed within five hours of initial access in around one case in ten. More strikingly, recent incident response reporting put the median time between initial access and hand-off to a second operator at twenty-two seconds.
That last figure describes an access broker economy rather than a single operator working through an intrusion at human pace. Somebody gains access, an automated handoff passes it to whoever bought it, and the person who eventually deploys the ransomware is not the person who got in. It matters operationally because the early stages look like commodity noise and the late stages look like a targeted attack, and they are the same intrusion.
Do not build a target time from those numbers. Build a stage preference from them: earlier detection is worth disproportionately more than faster response, and that is a statement about where the next engineering afternoon should go.
The distinction that matters
Those figures do not mean you must respond in seconds. Nobody does, and a program designed on that premise will produce automation nobody trusts and analysts who cannot keep up. What they mean is that the window between initial access and irreversible damage is measured in hours rather than in the weeks older breach statistics implied, and that older figures quoting months of average dwell time describe a different adversary from the one operating now.
They change which stage is worth investing in. If the adversary is through the handoff before your queue has refreshed, then detection coverage and automated containment carry more of the load than analyst speed can, and a plan that only makes triage faster is optimizing the stage with the least room left in it.
They also make the earlier stages disproportionately valuable. Detecting at initial access rather than at encryption is not a marginal improvement when the interval between them can be hours; it is the difference between an incident and an outage, and it is why the coverage question is about which stage a rule keys on rather than only whether a rule exists.
Track dwell time by incident type rather than as one figure. A ransomware precursor and an insider indicator have different expected detection windows, and averaging them produces a number that describes neither.
The split also protects the number from the case that distorts it most. One long-running investigation of a slow insider case will drag a quarterly mean into a range that makes the ransomware response look terrible, and a team that has just handled a precursor in forty minutes will reasonably feel the metric is lying about them. Three separate figures are harder to put on a slide and each of them means something.
Which Numbers Can Be Improved by Suppressing
A test to run on any pack, including your ownThe most useful thing you can do with a metrics pack is ask, of each number, whether suppressing more alerts would improve it. The answers are uncomfortable.
Microsoft Defender portal
Gives you the volume and closure figures most packs are assembled from, and Settings › Microsoft Defender XDR › Alert tuning tells you what was removed before those figures were counted. Read the second before quoting the first.
With both pages open, sort every measure in your pack into two columns and the shape of the problem appears immediately.
Figure EO0.8b. A measure that gets worse when you over-suppress is exactly the property worth having, and there is one of them.
Four of the five in the top block move together whenever a tuning rule is deployed, which means a pack made of them reports one action as five improvements. That is not dishonesty by whoever assembled it; it is a property of the measures.
Run the test on your own pack before you run it on anybody else's, and expect the answer to be uncomfortable. Most packs are assembled from what the tooling exports easily, and what the tooling exports easily is counts of things it processed, which is precisely the category that responds to filtering. Nobody chose the bad measures; they were the available ones.
Approach the conversation that follows accordingly. Arriving with the observation that four of the six numbers move together is a technical point about measurement; arriving with the implication that the team has been misleading people is a different meeting and a worse one. The measures were inherited in the same way the tuning rules were, and for the same reasons.
The useful version of the finding is not that the pack is wrong. It is that the pack cannot distinguish between two very different quarters, one where the team narrowed three rules and one where it suppressed three categories, and those have opposite consequences. A measure that cannot tell those apart is not measuring the thing its readers think it is.
The bottom block is harder to collect and it is the set that cannot be gamed by filtering. Two of them get worse when suppression goes too far, which is exactly the property you want in a metric.
Alerts per incident is the sharpest of them for that reason. It falls as suppression rises, because correlation is being fed a thinner stream, so it moves in the opposite direction to every number in the top block. A pack carrying both will show the tension rather than hiding it, which is the whole argument for including a measure that can embarrass you.
Set a floor for it rather than a target. A target invites gaming and a floor states a condition: below roughly one and a half alerts per incident, correlation is producing fragments rather than stories, and the tuning that produced the improvement elsewhere has gone too far. A floor is also easier to defend, because it is a statement about when to stop rather than about what to achieve.
Coverage, Again, and What It Is Against
The denominator problem in its most common formCoverage against a technique framework is the most quoted endpoint metric and the easiest to state without meaning, because it sounds specific and rests on two assumptions nobody states. EO0.5 established that it counts existence rather than function. There is a second problem underneath it, which is what the percentage is out of.
The enterprise technique matrix runs to a couple of hundred techniques and the exact count changes with each release, so a percentage computed last year against a different denominator is not comparable with this year's. More importantly, no estate needs all of them: techniques that require infrastructure you do not run are not gaps.
// The coverage figure that can be defended: rules that have demonstrably fired
AlertInfo
| where Timestamp > ago(180d)
| where DetectionSource == "Custom detection"
| summarize LastFired = max(Timestamp), Alerts = count() by Title
| extend Demonstrated = iff(Alerts > 0, "fired in period", "not demonstrated")
| sort by LastFired asc
That still does not prove a rule would fire on a real execution, because a rule can fire on benign activity and never on the technique it claims to cover. It is a floor rather than a ceiling: a rule that has never fired at all is certainly not demonstrated, and that alone will shrink most reported coverage figures.
Pair it with the validation runs from EO0.5 and you have the full picture. Fired in the period tells you the rule is alive; executed the technique and it fired tells you the rule works. The first is free and comes from a query, the second costs an afternoon per batch of rules, and a coverage claim built on both is one you can defend in a room.
Keep the validation dates in the same list as the rules rather than in a separate document. A coverage list where each row carries its own last-demonstrated date is self-auditing: rows age visibly, and a row with a date eighteen months old announces itself without anybody running a review. Split the two apart and the dates stop being maintained within a year.
The honest replacement is a named list rather than a percentage: technique, rule, date last demonstrated, data source it depends on. It fits on a page, every row is falsifiable, and it can say no.
Expect resistance to that, and expect it from a reasonable place. A percentage is comparable quarter to quarter and fits in a summary; a list of forty rows does not. The workable compromise is to report the count of demonstrated techniques rather than a proportion, keep the list underneath it, and never quote a denominator you have not chosen deliberately.
The denominator worth choosing is the techniques an adversary would plausibly use against this estate, which is a smaller and defensible set. Ten techniques covered out of a considered fifteen is a stronger statement than seventy-four per cent of everything, and it invites the right argument, which is about the fifteen.
Read a Metrics Pack
Six numbers, all improvingThe exercise below is the scenario at the top of this section. Read the notes under the table before the table itself, because the notes define what each number is measured on and that is where the findings are.
That ordering is a habit worth building for any pack, including ones you did not write. A table of numbers is unfalsifiable without its definitions, and the definitions are conventionally printed underneath in a smaller font because they are considered detail. They are the opposite of detail: the same six numbers with different definitions describe two different quarters.
The lines listed as not in the pack are the point of the exercise. Each of the five would change a decision, none of them is expensive to collect, and their absence is what makes the six numbers above them safe to report.
Notice also which numbers in the pack barely moved, because those are the honest ones. Incidents closed went from 1,190 to 1,204 across a quarter in which four other measures improved dramatically, which is what you would expect if the estate itself did not change. A number that stays flat while everything around it improves is usually the one measuring reality.
It is worth asking why that number was in the pack at all, since it is the one thing there that neither flatters nor guides. Incidents closed is a workload description: it says the function processed a similar amount of work to last quarter, which is useful for capacity planning and tells you nothing about security outcomes. Numbers like that belong in a resourcing conversation rather than a security one, and moving them there sharpens both.
The single-alert incident count is the other. It is not in the pack, it is 1,050 out of 1,204, and it is the direct consequence of the tuning that produced four of the six improvements. Every pack should contain at least one number that gets worse when the others get better, and that is the candidate here.
Capacity Is a Security Metric
And it is the one that predicts the othersEndpoint operations is people-constrained before it is tool-constrained, so the measure that predicts next quarter's numbers is whether the work fits in the hours available.
This matters more on an endpoint estate than elsewhere because of where the volume sits. Endpoint detection produces the large majority of alerts in a typical operation, so the endpoint queue is the queue, and a capacity problem anywhere in security operations usually turns out to be a capacity problem here.
Expected work is the alert volume that reaches an analyst multiplied by the time each one takes. Capacity is the analyst hours actually available for it after everything else. Both are estimates and both are worth computing anyway, because an approximate ratio changes decisions where an absent one does not. When expected work exceeds capacity persistently, the visible symptoms arrive in a predictable order: hunts stop first because they are the interruptible work, then investigation depth shortens, then tuning becomes suppression because it is faster, and then people leave. That order is worth memorizing, because each stage is a warning about the next, and the first two are visible in a metrics pack long before the fourth is visible in a resignation.
The distinction that matters
Expected work is alerts reaching an analyst multiplied by median handling time. At Northgate's Q3 that is 5,100 alerts at nine minutes, or 765 hours.
Capacity is the analyst hours actually available for the queue after everything else. Two analysts at roughly two thirds of their time is 624 hours.
The ratio, 1.23, is the number that predicts next quarter. Under 0.8 there is slack and hunting is possible. Between 0.8 and 1.0 it happens only if planned. Over 1.0 the deficit is being paid somewhere, and the payment is depth, tuning quality or people.
That ratio explains the hunt cadence in EO0.7 without anybody needing to be blamed. Six hunts planned and two executed is not a commitment problem at 1.23; it is arithmetic, and the honest response is to change the plan, the capacity or the queue rather than the expectation.
It is also the one number on this page that a manager can act on directly, which is why it is worth computing even though nobody asked for it. Every other measure here describes what happened; this one describes what is about to. A ratio above one predicts the depth and tuning quality that will show up in next quarter's incidents, months before those incidents occur.
Compute it honestly rather than optimistically. The fraction of an analyst's time actually available for the queue is rarely above two thirds once meetings, projects, on-call recovery and the rest of the job are removed, and using headcount as though it were hours is the single most common way this ratio gets reported as healthy.
Median handling time is the other input people guess at. Take it from the ticketing system rather than from an estimate, split by alert type, and expect the spread to be wide: a queue where most alerts close in two minutes and a few take ninety has a median that describes almost none of them. Where the spread is that wide, compute expected work from the distribution rather than the median, because the long tail is where the hours actually go.
Reporting Upward Without Lying
Translating a measure into a decision somebody else can takeThe measures above are for running the function. Presenting them is a separate skill, and the failure mode is not dishonesty but a number arriving with no decision attached to it.
The pattern that works is short. State the number, state what it means for exposure, state the decision it implies, and state what it would cost. Four sentences, and the reader can act on the fourth.
Microsoft Defender portal
And Reports › Device health give you the volume and health figures the pack usually starts from. Neither produces retrospective detection time or a validation record, which is why the numbers that survive scrutiny are the ones somebody has to assemble by hand.
Having assembled one, the question becomes how to say it, and the difference between two phrasings of the same finding is larger than it looks.
Figure EO0.8c. If you cannot name the decision a number implies, do not present it. That is a test of whether you understand your own measure.
Both are true and the same length. The second one names the decision, its cost and the consequence of not taking it, and it is the version that gets a fortnight allocated.
It also survives the follow-up question that ends most metric presentations, which is what would you do about it. The first version leaves that to the reader, who does not know the estate, and the answer they invent is rarely the one you wanted. The second has already answered it, which means the discussion is about whether two weeks is worth spending rather than about what the number means.
The corollary is worth stating plainly: if you cannot name the decision a number implies, do not present it. That is not a rule about communication, it is a test of whether you understand your own measure.
- Retrospective detection time, from first evidence in the logs, by incident type. The only version that describes the adversary.
- Detections demonstrated to fire, with a date against each. Replaces coverage as a percentage.
- Alerts per incident. Falls when suppression goes too far, which is why it is worth having.
- Open visibility gaps with owners. A list that accumulates and can be closed, rather than a number that fluctuates.
- Expected work over capacity. Predicts the other four next quarter better than any of them predicts itself.
A team that replaces its measures will produce a quarter where every number gets worse, because honest measurement of something previously measured loosely almost always does. Say so in writing before the first pack lands, with the old definition, the new one, and the same quarter computed both ways. Skip that and the improvement gets read as a decline.
Practice
Recompute one number honestly hands onOne incident of your own, recomputed honestly, is worth more than any argument in this section.
- Take one confirmed incident from the last quarter and find the earliest evidence of the activity in the logs, not the alert timestamp.
- Compute both detection times and put them side by side. The gap between them is what your current pack is not reporting.
- Run the suppression test down your own metrics pack, marking each number yes or no. Count how many are yes.
- Compute the capacity ratio for last quarter. Alerts reaching an analyst, median handling time, hours actually available.
- Write the five that survive on one page with the decision each would change beside it, and take that to whoever currently receives the pack.
The next section turns from how the work is measured to who it is being done against, and why an endpoint operator needs to understand technique rather than only tooling.