Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
0.7 The Investigator's Access
Introduction
Every lesson in this module so far has audited the record: what is captured, where it is kept, how it is protected. This one audits the people who will read it and act on it. An investigation needs read access to every account, from the first minute.
A response needs the ability to change things, to disable a key or attach a deny policy, without waiting for someone else to wake up. Both need a way in if the normal sign-in is unavailable or compromised. And the evidence they gather needs somewhere to go that the incident cannot reach. AWS's incident response guidance lists all four as preparation, to be in place before an incident.
Access is the part of readiness most often left to the incident itself to sort out. Organizations audit their logging, buy detection, and then discover during their first serious incident that the people who are supposed to respond cannot do anything without asking someone else, who is asleep, or on leave, or not sure they should.
Every hour spent arranging access is an hour the attacker's keys keep working. This lesson reads the record for the access Northgate's people actually used, audits the access they hold, and finds that Northgate's security team can read almost everything and change nothing.
The figure is the four things to have in place, in the order a response would reach for them. Read access lets an investigation start; change access lets a response contain; break-glass is the way in when everything else fails; and a forensic account is where the evidence goes. Each has a cost in risk, which is why each is decided carefully rather than granted broadly.
The audit in this lesson is about access granted on paper, and the record shows access used. The two together answer the question a readiness review most needs answered: when something goes wrong at night, who can do what, in which account, without asking anyone.
For Northgate the record of the period is the best evidence there is, because it shows every sign-in and every session the organization's people actually made.
Who Reads What
Permission sets by accountPeople reach Northgate's accounts through Identity Center permission sets, each a named bundle of permissions assigned to a group in chosen accounts, and the record shows which sets were used where. The query groups every permission set session's events by account and set.
SELECT recipientaccountid AS account, element_at(split(useridentity.arn, '/'), 2) AS permission_set_role,
count(*) AS events,
count(DISTINCT element_at(split(useridentity.arn, '/'), 3)) AS people
FROM cloudtrail_logs
WHERE useridentity.arn LIKE '%AWSReservedSSO_%'
GROUP BY 1, 2
ORDER BY 1, 2
Each permission set appears in an account as a role named for it, with a suffix Identity Center adds, and each session is named for the person who started it, which is how the query can count people.
The security team, three people, holds SecurityAudit in security, prod and dev, and used it in all three. AWS's SecurityAudit policy reads security configuration across services and changes nothing, which is right for everyday work. It is not assigned in the management account, where Identity Center, the organization's policies and the organization trail live.
An investigation into the two stolen Identity Center sessions would have needed to read the management account's activity and configuration, and the security team would have had to ask an administrator to do it for them.
Read access has a second gap the query does not show: the security team's SecurityAudit is a read of configuration, not of data. It can see that a bucket's policy changed, not read the objects in it, and it cannot take a snapshot or copy evidence anywhere.
For most of an investigation that is exactly right. For preservation, the steps of 10.6, a snapshot shared to another account, a memory capture run through Systems Manager, need more than SecurityAudit grants, which is one more reason the response permission set of section 5 should exist.
Management also holds the account where an investigator would check whether the organization's own controls changed: SCPs attached or detached, accounts added, trusted access for services turned on or off. A response that cannot read the management account cannot rule out that an attacker reached the top of the organization, which is the question every serious AWS incident eventually has to answer.
Automation, Not People
The cross-account roleNorthgate's security team also has a cross-account role, and the record shows what it is for. The query counts its assumptions and their callers.
SELECT recipientaccountid AS account, count(*) AS assumptions,
count(DISTINCT useridentity.arn) AS callers
FROM cloudtrail_logs
WHERE eventname = 'AssumeRole'
AND json_extract_scalar(requestparameters, '$.roleArn') LIKE '%SecurityAuditRole-CrossAccount'
GROUP BY 1
The role's assumptions are recorded twice, once in the security account where the call was made and once in prod where the role lives, which is why the query shows both. One caller, 37 times: the nightly audit function that reads bucket settings in prod.
The role exists only in prod, trusts the security account, and holds the same read-only SecurityAudit policy. It is a good example of a role built for one job, and it is not incident access: no person uses it, it reaches one account, and it can change nothing. Readiness reviews often find roles like this and count them as response access; the record shows otherwise.
Roles like this one should be in the review too, for the opposite reason: they are access nobody watches.
A read-only role assumed every night by one function is easy to forget, and an attacker who controlled the function's code could use the role's trust to read prod's security configuration from the security account. The audit records each such role, what assumes it, and whether that is still the only thing that does.
The cross-account role also shows a pattern worth copying for response. A role in each account, trusted by the security account, with exactly the permissions a job needs, is how automation reaches across an organization safely. A response role built the same way, assumable only by the security team's response permission set, would let containment reach every account from one place.
Northgate's one such role is a sound pattern; the gap is that nothing like it exists for response.
Emergency Access
Root and break-glassBreak-glass access is the most sensitive line in any access audit. It exists for the day normal sign-in fails: Identity Center unavailable, its identity source compromised, the administrators locked out. The query reads every use of the two emergency identities in the period.
SELECT eventtime, recipientaccountid AS account, sourceipaddress,
json_extract_scalar(additionaleventdata, '$.LoginTo') AS login_to,
json_extract_scalar(additionaleventdata, '$.MFAUsed') AS mfa
FROM cloudtrail_logs
WHERE useridentity.type = 'Root'
OR useridentity.username = 'breakglass-admin'
The root user is the other emergency identity: every AWS account has one, it can do anything in its account, and AWS recommends using it only for the few tasks that require it. The root user signed in once, on 25 April, to the billing console, from the office, with MFA, which is the kind of use AWS reserves root for.
The break-glass user, an IAM user in the management account with administrator access, a console password and MFA, was never used. Both are rare and explainable, which is right. Neither raised an alarm, which is not: if either had been used by someone who should not have, nobody would have known until they read the trail, and nobody reads the trail unprompted.
Identity Center itself is the dependency break-glass exists to bypass. If Identity Center's identity source is unavailable, every person at Northgate loses access at once, the security team included, at exactly the moment an attacker who caused the outage would want them locked out.
An IAM user in the management account, outside Identity Center, is the way back in. That is why it has administrator access, and why it is the most dangerous credential in the organization.
Session length is part of the access design too. Northgate's administrators get one-hour sessions and everyone else eight. A responder working an incident through the night signs in again every hour if they use administrator access, which is a small cost for limiting how long a stolen administrator session lasts; a response permission set should be set the same way.
Break-glass is also where many organizations discover their most surprising dependency: a password manager, a hardware token or a shared vault that itself depends on the identity provider that has just failed. Northgate's holders keep the credentials in a password manager, which is a question the drill should test.
Auditing the Access
Four gaps and two decoysThe access is summarized in one view for the audit, gathered from Identity Center, IAM and Organizations.
Mark the faults you would raise before reading on. The auditor shows the summary a readiness review would assemble from Identity Center's assignments, IAM's roles and users, and the organization's account list. Four gaps. There is no forensic account. There is no permission set that lets the security team contain anything.
SecurityAudit is not assigned in management. And the break-glass user's use raises no alarm. Two lines look like weaknesses and are not: the break-glass user has MFA, and administrator sessions last one hour, a deliberate limit on how long a stolen administrator session works.
The four gaps have different owners and different urgency. The response permission set and the management account assignment are decisions in Identity Center, made by its administrators. The break-glass alarm is a detection, the kind Module 9 builds, on two identities.
The forensic account is the largest piece of work, an account created and wired into the organization's sharing and storage, and the one most worth doing before rather than during an incident.
None of the four gaps is unusual, or a sign of carelessness. Most organizations grant their security team read access early, because it is low-risk and obviously useful, and postpone write access, break-glass alarms and forensic accounts because each raises a harder question about trust. A readiness review is the place to answer those questions deliberately, before an incident answers them by default.
Read Everything, Change Nothing
The missing response permission setThe second gap would have shaped every containment step in Module 10.
The record of the period shows no response at all, which 10.1 measured, and the access the security team held would have slowed any response that was attempted. Every step there, disabling terraform-ci's key, attaching AWSDenyAll, requiring IMDSv2, removing the lifecycle rule, required permission to change resources.
Northgate's security team has none. Those steps would have waited for one of three infrastructure administrators to be reached, briefed and willing, at whatever hour the incident was declared.
A response permission set fixes that: a set of exactly the containment actions, assigned to the security team in every account, with a short session, used only during a declared incident, and alarmed whenever it is used. It gives the responders the power to act and the organization a record that they did.
SecurityAudit: every account
IncidentResponse: every account
break-glass: alarmed
forensic account: existsRead permission set assignments, roles, break-glass and the account list.The gate's example is the night of 18 May, and the record lets the counterfactual be stated precisely. The backdoor user was created at 02:59 and GuardDuty's first finding about the escalation arrived at 03:07.
With Northgate's actual access, the security team could have read dev's record within minutes and contained nothing; with a response permission set, the backdoor key could have been disabled before the attacker's evasion attempts began at 03:22. The attempts were refused anyway, by the organization trail's design and an SCP; the point is that the response would not have had to rely on that.
A response permission set is a powerful thing to grant, and the design deliberately keeps it narrow. Its policy lists the containment actions and nothing else: deactivating keys, attaching AWS's deny policy, revoking role sessions, modifying instance metadata options, setting a function's concurrency, blocking public access, creating and sharing snapshots.
It does not grant deletion, which belongs to eradication and can wait for the people who own each resource. Its sessions are short. And because every use is a sign that an incident is under way, its use is alarmed and reviewed, like break-glass.
The objection is usually that the security team should not change production, and it deserves a direct answer. The answer is that in an incident someone must, quickly, and the alternative is waiting for an administrator who knows less about the incident than the team does. A narrow, alarmed permission set is a better control than a phone call at three in the morning.
Writing the response permission set's policy is also an exercise in knowing the response. Each action in it corresponds to a step in Module 10's lessons, and an action nobody can name a use for does not belong in it.
Somewhere for the Evidence
The forensic accountThe first gap shapes preservation, and it is the largest to close. Module 10's preservation plan copied snapshots, function code and logs into a locked evidence bucket, and it assumed somewhere to put them.
A forensic account
Security OUA forensic account is that somewhere: an account in the Security OU, accessible to the response team, into which snapshots are shared and copied, where volumes are restored for analysis, and where evidence is held under Object Lock.
Without one, a snapshot of app-prod-01 stays in prod, readable and deletable by whoever controls prod, and analysis happens beside the workloads under investigation. AWS's guidance recommends a dedicated account for exactly this, and Northgate's OU structure has room for it.
Sharing evidence into a forensic account has its own preparation, and it is easy to miss. EBS snapshots encrypted with the default AWS-managed key cannot be shared across accounts; they have to be copied with a customer-managed KMS key whose policy lets the forensic account use it.
An organization that discovers this during an incident spends its first hours re-encrypting snapshots. One that set up the key in advance shares a snapshot with a single call. The forensic account is only ready when the path into it has been tested.
The account also changes, fundamentally, who can see the evidence. In prod, anyone with administrator access, including an attacker who obtains it, could read or delete a snapshot. In a forensic account with its own access, only the response team can, and the evidence is out of reach of the incident it describes, the same principle that put the log archive in the security account.
Northgate's OU structure, Security and Workloads, gives the forensic account an obvious home under Security, where the organization's incident SCPs of 10.5 would not apply and where its access can be held as tightly as the log archive's.
Break-Glass, Alarmed
Rare, held and watchedThe fourth gap is the easiest to close and, in many organizations, the one most often left open.
Break-glass, as it should be
breakglass-adminBreak-glass works only if three things are true: it is available when everything else fails, it is held by few people, and every use is seen. Northgate's has the first two: an IAM user in the management account, independent of Identity Center, with MFA, its credentials held by two infrastructure leads. It lacks the third.
An alarm on any sign-in or call by breakglass-admin, sent to the security team and both holders, turns every use into a reviewed event, and turns a stolen copy of the credentials from a quiet backdoor into an immediate incident. The same alarm belongs on the root user of every account.
Holding break-glass credentials between two named people is the right design for an organization of Northgate's size, with one condition: each knows how to use them, and neither uses them alone without the other knowing.
The alarm makes the second condition enforceable. A yearly drill, in which one holder signs in, the alarm fires, the use is reviewed and the password is rotated, proves the whole path works and leaves the credentials fresh.
The root user deserves the same treatment in every account, not just management. AWS's centralized root access for organizations can remove member accounts' root credentials entirely, leaving tasks that need root to be performed from the management account on demand.
Where root credentials remain, they should be alarmed exactly as break-glass is, because a root sign-in nobody expected is one of the clearest signs of compromise an AWS account can show.
The alarm itself is simple to build directly from the trail's record: a detection on any event whose identity is breakglass-admin or a root user, run every few minutes, with no baseline and no threshold. Module 9 calls detections like this the most reliable kind, because the ordinary case is defined, here as never, rather than learned.
A Procedure for Access
Read, contain, break-glass, forensicsThe steps below audit the access any organization's responders would rely on. They are short because the decisions are few; what makes them hard is that each one is easy to postpone until it is needed.
useridentity.username = 'breakglass-admin'Each step's result also belongs in 0.8's assessment, which gathers every audit of this module into one picture.
Applied to Northgate, the four steps find read access in three accounts of four, no containment access at all, break-glass without an alarm, and no forensic account. Each is decided in the management account, by the people who administer Identity Center, and each should be tried once in a drill before an incident tries it for real.
The footer is why the drill matters: access that has never been used fails in ways nobody predicted.
Practice
Three questions about access in the record.
Next, 0.8 Assessing Investigation Readiness brings Module 0's audits together into one assessment, the one the course's Project asks you to make of an account of your own.