← Back to Blog

Your Time-to-Close Metric Ignores Every Incident Still Open

30 September 2026 Security Operations 9 min read
One month of incidents, and the half the metric can see 433 closed median time to close: 6 minutes 278 closed by the platform itself analysts: a median of 2.5 to 4 hours In the number 100 still open median age: 14 days; oldest: 30 days 57 never touched by anyone 20 of them High severity Not in the number Figures from one month of a test tenant's SecurityIncident table, 810 users.

Six minutes is an accurate median. It is the median of the incidents that finished.

A SOC lead opens the monthly dashboard and the headline is excellent: median time to close, six minutes. The same week, an analyst scrolling to the bottom of the queue finds a High severity incident from three weeks ago that nobody has opened. Both facts come from the same table. Only one of them made the dashboard, and that isn't a dashboard bug. It is how the standard query works, and it is why a SOC can report improving numbers for months while its queue quietly gets worse.

Nothing in this post needs a new tool. It needs three queries against the incident table you already have, and one change to how the result is reported.

Why the standard query cannot see the backlog

Microsoft's own guidance for measuring SOC efficiency in Sentinel takes each incident's latest version from the SecurityIncident table, subtracts CreatedTime from ClosedTime, and takes percentiles. It is a good query, and the latest-version step in it matters: the table writes a new row every time an incident changes, so counting rows counts changes, not incidents.

The problem is the subtraction. An incident that is still open has no ClosedTime. The difference is null, and percentile() skips nulls without a warning. So the metric is computed over closed incidents only, and every incident still waiting, including the oldest and the ones nobody has touched, silently drops out. The longer something waits, the less it counts, because it never finishes.

It also gets worse exactly when it matters. In a quiet month, most incidents close and the open tail is small, so the number is roughly honest. In a busy month, or a month with a real intrusion in it, the hard incidents stay open longest, the percentile loses them, and the reported close time can improve while the SOC falls behind.

That is survivorship bias in its purest form. A SOC that closes the easy incidents fast and leaves the hard ones sitting will report a better time to close than one that works the hard ones, and the report will be correct about every number it contains.

Run the standard measure and its missing half side by side:

// Table: SecurityIncident (Microsoft Sentinel, Log Analytics)
// Hypothesis: a time-to-close percentile describes only closed incidents,
// so it must be read beside the age of what is still open, or it misleads.
SecurityIncident
| where TimeGenerated > ago(30d)
| summarize arg_max(TimeGenerated, *) by IncidentNumber   // latest version of each incident
| extend HoursToClose = (ClosedTime - CreatedTime) / 1h,  // null while open
         OpenAgeHours = iff(Status != "Closed", (now() - CreatedTime) / 1h, real(null))
| summarize
    Closed           = countif(Status == "Closed"),
    CloseMedianH     = round(percentile(HoursToClose, 50), 1),
    CloseP90H        = round(percentile(HoursToClose, 90), 1),
    StillOpen        = countif(Status != "Closed"),
    OpenMedianAgeH   = round(percentile(OpenAgeHours, 50), 0),
    OpenOldestDays   = round(max(OpenAgeHours) / 24, 0)
// Read the two halves together. If StillOpen is more than a few percent of
// Closed, or OpenMedianAgeH is larger than CloseP90H, the close time is
// describing the easy work and the backlog is where the risk sits.

Against a month of our test tenant, the left half of that row says 433 closed with a median of six minutes and a 90th percentile of nineteen hours. The right half says 100 still open, waiting a median of 338 hours, the oldest for thirty days. The first number is what most dashboards show.

Separate the platform from the people

The six-minute median has a second problem, and it hides inside the closed half. Defender resolves a lot of incidents on its own: low-severity alerts, risks users cleared themselves by completing MFA, alerts that auto-resolved. Those closures take seconds, and in a busy tenant they can outnumber everything the analysts closed. Mixed into one median, they drag it toward zero and describe nobody's work.

ModifiedBy on the closing row says who closed each incident, so the fix is to group by it:

// Hypothesis: automatic closures by the platform, rules or agents take
// seconds and pull a blended median toward zero; each closer is its own
// population and needs its own number.
SecurityIncident
| where TimeGenerated > ago(30d)
| summarize arg_max(TimeGenerated, *) by IncidentNumber
| where Status == "Closed"
| extend HoursToClose = (ClosedTime - CreatedTime) / 1h
| summarize Closed = count(),
            MedianH = round(percentile(HoursToClose, 50), 1),
            P90H    = round(percentile(HoursToClose, 90), 1),
            NoVerdict = countif(isempty(Classification))
    by ClosedBy = ModifiedBy
| order by Closed desc
// Expect several populations: named analysts, automation rule names,
// agent identities, and the platform itself. Report the analysts' rows as
// the SOC's close time; report the rest as automation, with its own checks.

In the test tenant, 278 of the 433 closures came from the platform, 198 of them within a minute of the incident being created. The two analysts' medians were 2.5 and 4 hours, which is the figure that actually describes the team. The blended six minutes described neither.

The NoVerdict column is worth a glance while you're there. Automatic closures often carry no classification at all, which is fine for a low-severity alert the platform resolved and useless for learning anything from it. If your analysts' rows show blanks too, those are closures that will never tell a detection engineer whether the rule that raised them was right, and that is a separate fix worth making this month.

Numbers to hold your own report against. In the test month, the blended median time to close was 6 minutes, the analysts' own medians 2.5 and 4 hours, and the median age of the 100 open incidents 338 hours, about fourteen days. If your open incidents are older, at the median, than your closed incidents took at the 90th percentile, your backlog is the headline, whatever the close time says. And if more than about half of your closures are automatic, a single blended median is no longer describing your analysts at all.

Time to first touch has the same blind spot

The other number SOCs lead with is time to triage, which Microsoft's guidance computes from FirstModifiedTime. It has exactly the same survivorship problem, and a subtler one. An incident nobody has touched has no first modification, so it drops out of the percentile. And an incident an automation rule assigned in four seconds has a first modification that says nothing about whether a person ever opened it.

The honest version counts the tail directly, by severity:

// Hypothesis: incidents that were never picked up are invisible to
// time-to-triage; count them, and separate "assigned by automation" from
// "worked by a person" before believing the ownership column.
SecurityIncident
| where TimeGenerated > ago(30d)
| summarize arg_max(TimeGenerated, *) by IncidentNumber
| where Status != "Closed"
| extend AgeDays = (now() - CreatedTime) / 1d,
         Owner   = tostring(parse_json(Owner).assignedTo)
| summarize Open = count(),
            NeverTouched = countif(isempty(FirstModifiedTime)),
            OwnedButNew  = countif(Status == "New" and isnotempty(Owner)),
            OverSevenDays = countif(AgeDays > 7),
            OldestDays    = round(max(AgeDays), 0)
    by Severity
| order by Severity asc
// Any High row with NeverTouched > 0 is the first thing to fix, before any
// speed metric. OwnedButNew is work that looks assigned on a dashboard and
// has not started.

The test tenant returned twenty open High incidents, nineteen never touched, most of them single detections the platform raised on its own and nobody opened: directory attacks, a stolen session, ransomware preparation on a laptop. None of those nineteen affected the time-to-triage figure, because none had been triaged. The Low and Medium rows show the other gap: 42 incidents with an owner's name on them, still New.

The same read in Splunk

If your incidents reach Splunk from the Defender XDR incidents API, for example through the Splunk Add-on for Microsoft Security, the same two halves are one search. The API's incident fields are createdTime, lastUpdateTime, status (Active, Resolved or Redirected) and assignedTo, which is null when nobody owns it. It has no separate resolved time at the incident level, so this uses the last update on a resolved incident as the close, which is an overestimate if anyone edits incidents after closing them. Adjust the index and sourcetype to whatever your add-on writes.

index=security sourcetype="ms365:defender:incident"
| stats latest(status) as status, latest(severity) as severity,
        latest(assignedTo) as owner, earliest(createdTime) as created,
        latest(lastUpdateTime) as updated by incidentId
| eval created_e = strptime(substr(created, 1, 19), "%Y-%m-%dT%H:%M:%S")
| eval updated_e = strptime(substr(updated, 1, 19), "%Y-%m-%dT%H:%M:%S")
| eval hours_to_close = if(status="Resolved", (updated_e - created_e) / 3600, null())
| eval open_age_h = if(status="Active", (now() - created_e) / 3600, null())
| stats count(hours_to_close) as closed,
        median(hours_to_close) as close_median_h,
        perc90(hours_to_close) as close_p90_h,
        count(open_age_h) as still_open,
        median(open_age_h) as open_median_age_h,
        count(eval(status="Active" AND isnull(owner))) as open_unowned

The logic is identical: one row, closed on the left, open on the right, and the unowned count at the end, because an open incident with no owner is the purest form of the problem this post is about.

The sentence that keeps the report honest. A close-time percentile is a statement about finished work. Put the open count and the open age next to it every time it is reported, or the metric rewards leaving hard incidents open.

What an honest report looks like

The fix is not a better statistic. It is reporting the two halves together, every time, so that nobody can read one without the other. A report line that does that is short:

Closed by analysts        97   median 2.5 to 4 h, 90th percentile 18 to 30 h
Closed automatically     336   platform, agent and tuning rule, reported separately
Still open               100   median age 14 days, oldest 30 days
Never touched             57   19 of them High

Four lines, and every one of them came from the three queries above. The first line is the number the team can improve by working faster. The second is the number automation owns, and it deserves its own check, because a closure made in seconds by software is only as good as the rule or agent that made it. The third and fourth are the numbers the team improves by deciding what to work, and they are the ones a single close time hides completely.

Read the report the way a reviewer would. A falling close time with a growing open count means the team is getting faster at the easy work and leaving the hard work behind. A steady close time with a shrinking never-touched count means the queue is finally being read. Neither story is visible from the close time alone, and the second one is the one worth being able to tell.

Where to read more

  • Microsoft Learn, "Manage your SOC better with incident metrics in Microsoft Sentinel", which publishes the time-to-closure and time-to-triage queries this post extends.
  • Microsoft Learn, the SecurityIncident table reference, for CreatedTime, FirstModifiedTime, ClosedTime, ModifiedBy and Owner.
  • Microsoft Learn, "List incidents API in Microsoft Defender XDR", for the incident fields the Splunk search reads.

What to do this week

  1. Run the first query against your workspace for the last 30 days and write down all six numbers. If StillOpen is a surprise, the dashboard you report from has been showing half the picture.
  2. Run the second query and find your automatic closers. Rebuild your close-time figure from the analysts' rows only, and report automation as its own line with its own check.
  3. Run the third query and fix any High row with NeverTouched above zero today. Then decide, in writing, which severities the team will leave to age, instead of leaving it to chance.
  4. Add three columns to your monthly report: open count, open median age, and open unowned. Put them directly beside time to close, not on a later page.
  5. Compare the open median age with your close-time 90th percentile. If the first is larger, make the backlog the first line of the report next month, and measure whether it shrinks.

A fast time to close is good news only when you know what it left out. Count what finished and what is still waiting, in the same row, and the metric starts telling the truth.

Ridgeline Cyber Defence Written by security professionals. Published weekly on Tuesdays.

Related Articles

30 July 2026

The Automation That Reports Success and Does Nothing

A playbook that completes without acting reports green on every dashboard, because skipping a step counts as success. He

4 June 2026

KQL SigninLogs, The 10 Queries Every SOC Analyst Runs First

Ten KQL queries against SigninLogs that answer what SOC analysts actually ask during an identity investigation, copy-pas

3 May 2026

Five KQL Threat Hunts Every M365 SOC Should Run This Month

Your detection rules cover known patterns. These five KQL hunts find the attacker activity that bypasses every analytics