Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
0.8 Assessing Investigation Readiness
Introduction
Module 0 has audited Northgate one part at a time, lesson by lesson: the trail, what it records, where the record is kept, how detection works, and what access responders hold.
An assessment puts those audits together into one statement of how ready the organization is to be investigated, written for the people who decide what to fix. It has to be accurate in both directions, about strengths and about gaps. An assessment that overstates readiness is worse than none, because leadership will rely on it and discover its errors during the incident it was meant to prepare for.
The assessment is also the module's bridge to the rest of the course. Every gap it names is one the investigations of Modules 4 to 8 run into, the detections of Module 9 are built around, or the response of Module 10 works around, and 10.8's review closes.
Reading the assessment first and the course second shows each gap's cost in a real incident; reading the course first and the assessment second shows why each line matters. This lesson measures the five areas from the record, audits a colleague's draft assessment that gets four of them wrong, orders Northgate's gaps by urgency, and sets out the procedure, which is also the first part of the course's Project.
The figure is the assessment's structure, part by part, and each part draws on one or two of the module's lessons. The verdict is the part leadership reads first: the gaps, who owns each, and the order in which to close them.
Each query in this lesson is one an assessor can run, almost unchanged, against their own organization's record, with only the account numbers and bucket names changed. Together they turn the module's audits into numbers that can be compared from one assessment to the next.
The figure's five parts are also the order in which to write the assessment, because each depends on the one before: there is no point assessing how long a record is kept until it is clear what it records.
What Is Recorded
Coverage in one rowThe first part of the assessment is what the record covers, and a single query over the whole record states it.
SELECT count(DISTINCT recipientaccountid) AS accounts,
count(DISTINCT awsregion) AS regions,
count_if(eventcategory = 'Data') AS data_events,
count(DISTINCT CASE WHEN eventcategory = 'Data'
THEN json_extract_scalar(requestparameters, '$.bucketName') END) AS buckets_with_data_events
FROM cloudtrail_logs
The query counts the accounts, Regions, data events and buckets with data events across the whole record. Four accounts and seven Regions are 0.3's organization trail doing exactly what it should. Coverage of every account and Region is the foundation of everything else in the assessment, and Northgate has it.
The assessment says so plainly, because a reader needs to know what is solid before reading what is not. Three buckets with data events, and none of the record's events from a Lambda invocation, are 0.4's edges. The edges are where the assessment's language matters most.
Saying that object access is recorded for three named buckets, and naming the bucket that is not, gives leadership a decision to make; saying that data events are enabled gives them false comfort. An assessment states both halves in the same sentence, because a reader who hears only the first will assume the second.
The coverage line should also say what changed during the period, if anything did. Northgate's trail and selectors did not change between 14 April and 21 May, which Module 8 confirmed. An assessment that covers a period in which coverage changed has to say when, because a gap that existed for a week is a week no investigation can see into.
Network evidence belongs in the coverage line as well. The prod VPC's flow logs, the resolver's DNS logs and the load balancer's access logs are all recorded, and each answered a question in the course's investigations that CloudTrail could not. The assessment lists them with the trail, because an organization that records API calls and nothing else can answer only part of what an investigation will ask.
Coverage, in short, is strong where it is broad and thin where it is chosen, and the assessment's first paragraph should say exactly that.
What Is Kept
Findings and their deadlineRetention is easiest to understand as a date on a calendar, so the assessment converts it into one.
SELECT 'guardduty findings' AS item, count(*) AS n,
min(date_add('day', 90, from_iso8601_timestamp(createdat))) AS first_lost
FROM guardduty_findings
The query takes the period's findings and the date the first of them will leave GuardDuty. Eleven findings, the first gone on 9 August unless exported. Findings are not the only deadline.
The log archive's retention is longer, 400 days, but its first 90 days are the only ones readable without a restore. The assessment states each source's retention and the date it matters, so that a gap becomes a deadline.
Integrity belongs in the same part of the assessment as retention, deliberately, because both are about whether the record will be there and be trusted when needed. Northgate's archive is validated, which lets an investigation prove its record unaltered, and is not locked, which means it could be altered. The assessment states both, and names the role that could do the altering.
The archive's readability belongs here too: everything older than 90 days needs a restore before Athena can read it. For the assessment that becomes a plain sentence, that an investigation reaching back more than three months will wait hours for its first answer, and a recommendation to change the storage class.
Each date the assessment states is also a test of the organization's process: if a finding's expiry passes without an export being set up, the assessment's recommendation was not acted on, and the next assessment should say so.
What Can Be Restored
Configuration historyThe history an investigation can restore from is a short line in the assessment and a long problem during an incident.
SELECT count(*) AS config_change_items,
count(DISTINCT ci.awsaccountid) AS accounts_with_history
FROM awsconfig
CROSS JOIN UNNEST(configurationitems) AS t(ci)
WHERE configsnapshotid IS NULL
The query counts Config's change items, across the whole period, and the accounts they come from. One account, dev, nineteen items, covering instances and three IAM types. The assessment states which resource types have history in which accounts, because the answer decides what a future response can restore exactly and what it must reconstruct.
Prod, where the backups were stolen, has no history for its buckets or its IAM, and Module 10 showed the consequence: the backup bucket had to be restored from its owner's written intent.
Detection has its own line in the assessment, which this lesson's queries do not need to repeat: GuardDuty and Security Hub run for every account from the security account, two protections that match Northgate's workloads are off, and no finding has an owner.
The last point is the one 10.8 found mattered most, and an assessment that leaves it out because it is not a configuration setting leaves out its most important finding.
History can also come from the trail itself, more slowly. Where Config has nothing, an investigator can reconstruct a resource's past by reading every call that changed it, back to its creation, if the trail still holds that far. For a bucket created years ago that is rarely possible, which is why the assessment treats Config's gaps as gaps, not as inconveniences.
Ownership is the same in the assessment as in the review: every gap needs one, or it stays a finding. For the history gap the owner is whoever administers prod's Config recorder, and the change is a recording group that adds IAM roles and policies and S3 buckets; for the readability gap it is whoever owns the log archive's lifecycle.
Who Can Act
Access in the recordThe access part of the assessment is stated from what the team actually used in the period, not from what it is assumed to have.
SELECT count_if(useridentity.arn LIKE '%AWSReservedSSO_SecurityAudit%') AS security_team_reads,
count(DISTINCT CASE WHEN useridentity.arn LIKE '%AWSReservedSSO_SecurityAudit%' THEN recipientaccountid END) AS accounts_read,
count_if(useridentity.type = 'Root' OR useridentity.username = 'breakglass-admin') AS emergency_uses
FROM cloudtrail_logs
The query counts the security team's SecurityAudit events, the accounts they fall in, and every use of the root user or break-glass. Stating access from the record has an advantage over stating it from configuration: it shows what is actually used.
A permission set nobody has assumed is access in name only, and a role assumed only by automation is not a person's access at all. 525 read-only events in three accounts; one emergency sign-in, the root user's billing visit. The team can read most of the organization and change none of it, which is the single most important line in Northgate's assessment.
Break-glass and the forensic account complete the access picture. Northgate's break-glass user exists, has MFA and has never been used, and its use would raise no alarm. There is no forensic account. Neither appears in the record's numbers, which is exactly why the assessment has to state them: absence leaves no event to count.
The four parts measured so far are the evidence. The verdict comes after the draft is corrected, because a verdict built on an overstated draft orders the wrong gaps.
Reading the access line, leadership should come away knowing three things: who can investigate, in which accounts; who can stop an attacker, and how long that would take to arrange; and what happens if normal sign-in is lost. Northgate's honest answers are the security team, in three accounts; only the administrators, after a call; and break-glass, unalarmed.
Those three answers are the assessment's access line, nearly word for word.
Auditing a Draft
Four overstatementsAssessments are usually written by someone who did some of the audits and assumed the rest, often under time pressure, and often with the best intentions. The auditor below is a colleague's draft for leadership.
Mark the lines you would challenge before reading on; the explanation compares each with the lesson that audited it. Four lines overstate Northgate's readiness, and each would mislead leadership in a specific way: that every object and function is recorded, that the archive cannot be altered, that findings are permanent, and that the security team can act.
Two lines are correct and should stay. The four errors share a single cause: each describes what the writer assumed a well-configured organization would have, not what Northgate has. The gate below is the test every line should pass.
configuration read
query run
gap statedCheck each line against the audit that should support it.Applied to the draft, the gate keeps two lines and rewrites four, each from the lesson that audited it. The corrected assessment is less flattering and far more useful: it tells leadership where the money and the effort should go.
The draft is a realistic and common kind of error. Its writer knew what a well-run organization should have and assumed Northgate had it, because nothing obvious contradicted the assumption. Every line sounded right. That is why each line needs its evidence beside it, a configuration read or a query run: the assessment's reader cannot tell a checked line from an assumed one, and the writer often cannot either.
The two correct lines in the draft matter too. An auditor who challenges everything teaches leadership that the assessment is unreliable as a whole; one who confirms what is right and corrects what is wrong teaches them which parts to trust. The trail and the delegated detection are genuine strengths, and the corrected assessment should lead with them.
The corrected lines are the ones each audit would have written: three named buckets and two gaps, a versioned but unlocked archive with a role that can delete from it, findings kept ninety days and not exported, and a security team that can read three accounts and contain nothing. Each is less comfortable than the draft and each gives leadership something to decide.
The Gaps in Order
What to fix firstThe verdict orders the gaps by how much each would cost an investigation or a response, and how soon. The record below is Northgate's list, gathered from the five lessons that found each gap.
Northgate's readiness gaps, in order
From 0.3 to 0.7Response access comes first in Northgate's order, because it decides whether the first hour of an incident is spent acting or waiting. The log archive comes second, because a role that can delete the record is a present risk. Findings come third, with a date.
Coverage and readability matter on the day of a particular kind of investigation, and the housekeeping items are cheap and can be done alongside. Each gap in the list has a named owner, the team that administers what has to change, and each is the subject of an action in 10.8's review.
Ordering the gaps is a judgment, not a calculation, and the assessment should show its reasoning so that leadership can disagree with it on the merits. Northgate's order puts response access first because the record shows a period in which detection worked and nothing happened, and access is one reason nothing could happen quickly.
An organization with a different history might order differently: one that had lost evidence to a deleted log file would put the archive first, and one whose incidents were in unrecorded buckets would put coverage first.
Each gap also needs a cost, even a rough one, because leadership weighs cost against risk. Most of Northgate's are configuration changes with small ongoing costs: Object Lock, an export bucket, a permission set, a Config recorder change.
Runtime Monitoring and data events for busy resources carry per-event charges worth estimating. The forensic account is mostly effort. An assessment that lists gaps without costs invites the answer that everything is too expensive; one with costs usually shows that most of it is not.
The list also connects directly to the end of the course. 10.8's review assigned actions for the causes of the period's attacks; the gaps here are the readiness causes behind them, and the two lists together are the complete set of changes Northgate needs. An organization that runs this assessment before an incident is doing 10.8's work in advance, at a fraction of the cost.
Ordering also tells leadership what can wait. Readability and some coverage gaps cost little to live with until a particular kind of investigation needs them, and an organization with limited time is right to close response access and the archive first. An assessment that ranks everything as urgent gives no guidance at all.
What the Assessment Is Not
A snapshot, with limitsAn assessment states readiness on the day it was made, from the configuration and record that existed then. The record below sets out what it can and cannot claim.
What a readiness assessment can claim
Its scopeIt is not a statement that no attack will succeed; Northgate's record shows attacks succeeding against an organization with better-than-average readiness. It is not a measure of detection quality; Module 9 measures that by testing detections against incidents. And it is not permanent: every new account, bucket, function or person changes it.
What it does is tell an organization, before an incident, which questions an investigation could answer and which it could not, and what it would take to close the difference. That is the decision readiness exists to inform.
The limits matter most when the assessment is used, as assessments often are, for something it cannot support. A readiness assessment is sometimes asked to answer whether the organization is secure, or whether an incident could have been prevented. It answers neither. It answers whether the organization could investigate and respond to one, which is a narrower and more useful question, and the one this module set out to answer.
Northgate's own period illustrates the difference clearly and in detail. Its readiness, measured by this module, was better than many organizations': one trail for everything, an archive out of the workload accounts' reach, detection in every account. It was attacked six times in five weeks anyway, and its record let every attack be investigated in detail.
Readiness made the investigations possible; it did not prevent the attacks, and it was never meant to.
The record above draws the line. Everything on the can-claim side is something the module's audits measured; everything on the other side needs evidence the assessment does not collect. Keeping to the line is what lets leadership trust the parts the assessment does claim.
A Procedure for Assessment
Five parts, every line checkableThe steps below produce a readiness assessment for any AWS organization.
The procedure's last step is the one most often skipped because it is the least comfortable: turning findings into a verdict with an order and owners means telling people their area has a gap and asking them to close it by a date. An assessment that stops at findings is a report; one that ends with owned, ordered actions is a plan.
The assessment is also short, which is a feature. Northgate's fits on two pages: five parts of a paragraph or two each, every line with its evidence beside it, and the ordered gaps at the end. Length is not a measure of thoroughness, and a long assessment is more likely to be skimmed by the people who have to act on it.
Applied to Northgate, the procedure produces an assessment with a strong foundation, an organization trail and delegated detection, and six gaps led by response access and the log archive. Every line is backed by a configuration or a query from this module. The footer is why it is rerun: the organization changes faster than any single assessment can describe.
Practice
Three questions about the assessment's evidence.
Module 0 ends here. Module 1, Querying CloudTrail with Athena, begins querying the record this module decided to keep.