In this section

How to Measure Identity Detection Coverage

Module 0

Introduction

Every module in this course produces a number somebody will want reported, and most identity security numbers are constructed so that they can only improve. You will finish this section able to say what makes a coverage figure honest, why the denominator carries the whole argument, and which four numbers are worth putting in front of somebody who will act on them.

Scenario

A quarterly security report shows identity coverage at 94 per cent, up from 91. The figure is the proportion of users covered by at least one Conditional Access policy. Nobody has misrepresented anything: the calculation is correct, the data is current, and the trend is real. The 6 per cent excluded contains every break-glass account, four service accounts with mailbox access, the CEO, and eleven contractors added during a project. The number went up because 90 new starters were onboarded into the covered population, and none of the exclusions was reviewed.

01

A number that can only go up

Which is how you know it is the wrong one

The 94 per cent is arithmetically correct and it is not a measure of anything anybody cares about. Its numerator grows every time somebody joins the organization and its denominator grows at the same rate, so the ratio drifts upward as the covered population expands, entirely independently of whether the estate got safer.

That property is the tell. A metric that improves when nothing changes is measuring activity rather than outcome, and identity security is full of them: MFA registration rates, policies deployed, rules enabled, alerts triaged, training completed. Every one of those can rise for a whole year in an estate that is becoming less defensible.

The scenario's figure has a specific mechanism worth naming. Its numerator is users covered and its denominator is all users, and both grow with headcount, but they do not grow at the same rate: new starters are onboarded into the standard groups the policies target, so every joiner lands in the numerator automatically while exclusions are added one at a time by exception. A growing organization therefore reports improving coverage as a side effect of hiring.

Reverse it and the mechanism runs backwards: a quarter of redundancies shows coverage falling, because the covered population shrinks while the exception list does not.

These are not bad-faith numbers. They are easy to collect, easy to explain and easy to trend, which is what a reporting cycle rewards. Nobody chose them to mislead; somebody needed a slide in a fortnight.

A harder version is a number that can fall only for reasons nobody controls. Alert volume is the usual example: a quieter quarter reads as improvement and is as likely to mean a rule stopped firing or a table stopped arriving. Movement either way is uninformative unless you know which.

The test that separates a good metric from a bad one is not whether it can fall. It is whether you can name, in advance, the specific thing that would make it fall, and whether that thing is something you would want to know about.

Keep this

Ask what would make this number go down. If the honest answer is "nothing we would notice", the number is not measuring the estate, it is measuring the size of the organization or the effort of the team.

That question is the whole method, and it is worth applying deliberately rather than by instinct, because the numbers that fail it are the ones that feel most reassuring to report. Applied to the four figures a typical identity report carries, it sorts them into two groups of two, and the sorting is not the one most people expect.

02

The test that separates the two kinds

Name the thing that would make it fall

Both problems above have one test in common, and applying it to four familiar numbers sorts them cleanly.

The test is to name, in advance and specifically, the thing that would make the number fall. Not whether it can fall in principle, which almost any number can, but what event in the estate would move it downward and whether that event is something you would want to hear about.

It also has to be asked of numbers you produced yourself and are pleased with, which is the harder case: a metric built to demonstrate progress passes every check except this one.

MFA registration fails immediately. Nothing plausible reduces it: people do not un-register methods, and the only mechanism that lowers the figure is somebody joining without one, which onboarding prevents. The number is a record of a project that finished.

Alerts triaged fails differently, and more dangerously, because it looks like it passes. It can fall, and the most likely cause is a rule quietly dying or a table ceasing to arrive, so a falling value is evidence of a problem being reported as an improvement. Per section 0.1 that is the exact failure the estate there could not see.

The two below the line pass because their failure directions are meaningful. Identities with no detection rises when somebody creates a service principal, which is a real event you want to know about and one that happens without anybody telling you. True positives per rule falls when a rule stops catching things, which is either the threat changing shape or the rule breaking, and both of those warrant somebody looking.

Two kinds of number, and the test that separates them the metric what makes it fall verdict MFA registration rate nothing you would notice effort, not outcome alerts triaged a rule silently dying falls for the wrong reason identities with no detection somebody doing the work reportable true positives per rule, 90d a rule that stopped catching reportable, and nobody keeps it The test is not whether a number can fall. It is whether you can name in advance what would make it.

Figure 0.14. The top two rise for a year in an estate getting less defensible. The bottom two cannot.

Effort metrics are not worthless, and they are not what a security report is for. A team that triaged 1,900 alerts did real work; the mistake is presenting that as evidence the estate is defended, which is a different claim needing different numbers. The next section is about what the denominator on those numbers should be.

03

The denominator is the argument

What 94 per cent left out

Every coverage figure is a fraction and almost all the meaning sits in what the bottom half includes. The scenario's denominator is users, and per section 0.6 that excludes 40 per cent of the identities in the tenant before the calculation begins.

The same estate, measured four ways Every figure below is arithmetically correct. Only the last one is a coverage number. users covered by a policy 94% ...including service principals 55% ...covered by an ENABLED policy 41% ...with a detection watching it 14% Each row adds one honest condition to the denominator. Nothing about the estate changed between them.

Figure 0.13. The reported figure and the useful one differ by eighty points, and both describe the same tenant on the same day.

Read the figure at the end of each bar rather than the bar itself. The estate is identical in all four rows: same tenant, same day, same policies, same identities. The only thing changing is what the fraction is being taken over, and the reported figure and the useful one differ by eighty points.

Each row adds one condition somebody would agree with if asked. Should service principals count as identities that need covering? Obviously, and per section 0.6 they are 40 per cent of the tenant. Should a policy sitting in report-only count as covering anything? Obviously not, and per section 0.7 four of them had been there for fourteen months. Should coverage mean somebody would find out if the control failed rather than that a control exists? That is the question this course exists to ask, and it is the one that takes the figure from 41 to 14.

The last row is the one that generates argument, because it changes the meaning of the word coverage from a property of the control to a property of the estate. Covered by a policy means a rule would evaluate the sign-in. Watched by a detection means somebody would find out if the policy did not do what it was supposed to. Section 0.1's estate had the first for almost everybody and the second for almost nobody, and lost nineteen days in the gap.

None of those adjustments is a trick. Each is the honest answer to a question the original figure never posed, and the aggregate is the difference between a number that reassures and a number that directs work.

04

Exclusions are where coverage actually lives

And they are always the interesting accounts

The 6 per cent in the scenario is not a random sample of the population. Exclusions accumulate for reasons, and the reasons select for exactly the accounts worth protecting.

Entra Admin Center

ProtectionConditional AccessInsights and reporting
The workbook here reports on policy impact over a time range, which makes it an events view rather than a state one. Per section 0.8 it tells you what policies did to sign-ins that happened, and not what they would do to a population that has not signed in.

Break-glass accounts are excluded by design and correctly so, because a policy that locks out every administrator must not lock out the recovery path. Service accounts are excluded because they broke when a policy applied. Executives are excluded because a policy caused friction and somebody senior asked. Contractors are excluded because a project needed access on a deadline and the exception was never revisited.

That is four different mechanisms producing one list, and only the first is a decision anybody would defend today. The list of exclusions is a better description of an estate's real posture than the coverage figure, and it is shorter, and nobody reports it.

There is a specific pattern worth watching for in that list, which is the same account excluded from several policies. A break-glass account excluded from everything is the design working. A service account excluded from three policies is an account that broke three times and was exempted three times rather than fixed once, and it usually holds more access than anybody realizes.

Connect-MgGraph -Scopes "Policy.Read.All","Directory.Read.All"

Get-MgIdentityConditionalAccessPolicy -All |
  Where-Object { $_.Conditions.Users.ExcludeUsers -or $_.Conditions.Users.ExcludeGroups } |
  ForEach-Object {
    $p = $_
    $p.Conditions.Users.ExcludeUsers | ForEach-Object {
      [pscustomobject]@{ Policy = $p.DisplayName; State = $p.State
                         Excluded = (Get-MgUser -UserId $_).UserPrincipalName }
    }
  } | Sort-Object Excluded

# Policy                        State    Excluded
# ------                        -----    --------
# Require MFA, all users        enabled  admin.break1@northgateeng.com
# Require MFA, all users        enabled  r.okafor@northgateeng.com
# Require compliant device      enabled  svc-reporting@northgateeng.com
# Block legacy auth             enabled  svc-reporting@northgateeng.com
# ...

Two things about that command are worth noting. It reports the policy state alongside each exclusion, because an exclusion from a report-only policy is not an exclusion from anything and counting it as one overstates the problem. And it resolves the object ID to a UPN, which sounds trivial and is the difference between a list somebody reads and a list of GUIDs somebody closes.

Run that once and the output is usually shorter than expected and worse than expected. Per section 0.7 nothing prompts a review of an exclusion, because an exclusion produces no events, no alerts and no errors: it is the absence of something happening.

05

The four numbers worth reporting

All of which can go down

Replacing a bad metric with a good one is harder than criticizing it, so here are four that survive the test in section 1 and that this course teaches you to produce.

Identities with no detection watching them. Per section 0.6, counted across the whole tenant rather than across users. It goes up when somebody creates a service principal and down only when somebody does work.

Rules that cannot fire. Per section 0.1, the count of enabled rules whose table is not ingested. It should be zero, it usually is not, and the number is unarguable in a way most security metrics are not: there is no interpretation under which a rule querying an empty table is doing anything.

True positives per rule, over ninety days. Per section 0.3, the only figure that distinguishes a rule watching something rare from a rule watching nothing. It is the number nobody keeps, because it is the one that can fall.

Controls tested since the last report. Per section 0.7, break-glass, policy restore, and anything else whose failure mode is silence. Expressed as a count with dates rather than a percentage, because a percentage hides which one has not been tested and the identity of the untested control is the whole information.

Notice what those four have in common structurally. Each has a denominator that is a real population rather than a convenient one, each has a definition that fits in a sentence, and each can be produced from the surfaces section 0.8 described without any new tooling. Three of the four are Graph queries and the fourth is a date somebody writes down.

None of them requires a platform, a dashboard product or a maturity model. That matters because the usual response to bad metrics is to buy something that produces more of them.

Keep this

A number worth reporting is one that can embarrass you. All four of those can move in the wrong direction without anybody doing anything wrong, which is precisely what makes them useful: they describe the estate rather than the effort.

Each of the four is produced by a module of this course, which is deliberate: id01 gives you the first, id04 the second and third, and id06 the fourth. By the end you can report all four from your own tenant.

06

Reading the report

Find the three claims that do not hold

Below is the quarterly report from the scenario as it was presented, with the underlying data beside each claim. Nothing in it is fabricated, and nobody who wrote it was trying to mislead anybody, which is what makes it worth reading carefully rather than dismissing.

The tool below does the same job on your own estate. It takes the output of the commands from sections 0.2 and 0.6 rather than asking you to rate yourself.

Nothing in it is weighted or benchmarked. Where your input cannot support a figure it says so rather than estimating, which is what this section asks of a report.

The exercise is not to find an error, because there is not one. Three of its claims fail that test and they fail in different ways. One is a figure whose denominator excludes most of the identities in the tenant. One is a trend produced by a mechanism unrelated to security. And one is a statement about what was caught that is silent about what was looked for.

It is to identify which claims are true statements that would lead a reader to a false conclusion, which is a different reading skill and the one worth having in front of a report you did not write.

**Work each claim against the data beside it and ask what a reasonable reader would conclude, rather than whether the sentence is accurate. Those are different questions and the gap between them is where a report misleads without containing a falsehood.

The second claim is the one most reports contain. A statement that the estate detected and responded to a number of identity threats this quarter is true, verifiable, and silent about the threats it would not have detected, which is the number nobody has. Per section 0.4 an estate can block 41,200 password attempts and lose 69 days to three attacks it was never watching for.

The hardest of the three is the trend.** Coverage rising from 91 to 94 is a real change in a real number, and the reason for it, ninety new starters entering the covered population, has nothing to do with security work. A reader has no way to know that from the slide, and neither did the person who made it.

07

Reporting to somebody who will act

And what happens when the number falls

There is a practical objection to everything above, and it deserves an answer rather than a dismissal: honest numbers are harder to report, and a figure that gets worse invites questions the team may not want.

The answer is to introduce the falling number alongside the reason it falls. A count of identities with no detection watching them that rises from 340 to 380 is not a failure of the security team, it is 40 new service principals somebody else created, and presented that way it is an argument for a governance process rather than an admission.

It is worth saying that this is easier to do at the start than in year three. A team introducing an honest number for the first time can present it as a new measurement rather than as a deterioration, and the first value it takes is a baseline rather than a failure. A team that has been reporting 94 per cent for two years and switches to 14 has a harder conversation, and the temptation at that point is to keep reporting 94.

What does not work is switching metrics when the current one turns uncomfortable. A report that changes its definition of coverage between quarters has destroyed the only thing a metric is for, which is comparison against itself.

Pick numbers you are willing to report when they are bad, define them once, write the definition down next to the figure, and keep it. The definition matters more than the number, because a figure without one cannot be checked, and per section 0.8 the same estate honestly produces 94 per cent and 14 per cent depending on what the denominator was.

08

Practice

Recalculate one of your own figures

Thirty minutes against a number your organization already reports.

Do this Find the denominator
  1. Take one identity coverage figure your organization reports and write down its denominator exactly. If nobody can state it precisely, that is the finding and you can stop there.
  2. Recalculate it including service principals, per section 0.6. Note both numbers and the gap between them.
  3. Recalculate again counting only enabled policies, excluding anything in report-only.
  4. Run the exclusions command above and read the list. Ask of each entry whether anybody would defend it today, and note how many you cannot answer for.
  5. Write the definition down beside the figure. That sentence is worth more than the number and it is the thing that makes next quarter's version comparable.

What you should end up with: one figure expressed three ways, a list of exclusions nobody has reviewed, and a written definition. The gap between the first and third numbers is the honest measure of what the reporting was hiding, without anybody having lied.

The last section of this module covers what you need to follow the rest of the course in a tenant of your own.

// reasoning-review IS NAMED HERE BECAUSE IT EMITS ITS rc-code AT RUNTIME. // This condition scans the SERVER-RENDERED content, and reasoning-review.js builds its // artifact block after fetch, so the page contains no rc-code at the moment this runs and // code-chrome never loaded. Deployed 2026-08-26 with correct markup, transparent background // and no chrome, because the class the loader looks for did not exist yet. Any future // component that writes rc-code from script has to be named here too.