Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
0.5 Retention, Integrity and the Log Archive
Introduction
A trail that records everything is only half of readiness. The other half is what happens to the record after it is delivered: how long it is kept, whether an investigation can read it on the day it is needed, and whether anyone could have changed it in between.
Those properties belong to the bucket the trail delivers to, and to the account that bucket lives in. AWS's guidance, and Northgate's design, put that bucket in a dedicated log archive account, out of reach of the workload accounts it records. Readiness for the archive is decided in a handful of settings, all in one account.
This lesson audits Northgate's log archive: the volume it holds, the account it sits in, how long it keeps the record and in what storage class, and who can change it. It finds a bucket that is well placed and versioned, and four settings that would each weaken an investigation that needed the record most.
The figure is the three properties an archive needs, and the account that protects all three. A record kept too briefly fails a late investigation; one kept in an unreadable class delays it by hours; one that can be altered cannot be relied on in front of anyone who doubts it.
Northgate's archive is also the evidence every module of this course reads, so its gaps are not abstract. If the backups theft of 16 May had been discovered in September rather than days later, the investigation's first queries would have waited for a restore; if anyone with SecurityAdmin had deleted the files for 18 May, validation would have shown the gap and nothing could have filled it.
What the Archive Holds
Volume and retention costRetention decisions are almost always framed as cost decisions, so it is worth knowing what the record costs to keep. The query measures the trail's daily volume over the period.
SELECT count(*) AS days, min(events) AS quietest_day, max(events) AS busiest_day,
approx_percentile(events, 0.5) AS median_day
FROM (
SELECT substr(eventtime, 1, 10) AS day, count(*) AS events
FROM cloudtrail_logs
GROUP BY 1)
The query counts events per day and summarizes the days. Northgate's trail records a median of 734 events a day across four accounts, small for an organization of its size because its workloads are steady and few.
Even at the busiest day's rate, a year of management and data events is a few hundred thousand records, a few hundred megabytes of compressed files. Keeping the record for years costs less than most of the tools that read it. Retention is limited by decisions, not by money, and the decisions are written in the lifecycle rule this lesson audits.
Data events change the arithmetic for busy buckets, and network activity events can too; an organization that records millions of object reads a day has a real storage bill.
Even then, the cost of keeping is small beside the cost of the question an investigation cannot answer because the record was deleted. The right retention is the longest period over which an investigation or an obligation might need to look back, and many organizations find that is longer than they assumed.
The log archive account is a pattern AWS recommends for multi-account organizations: one account whose only job is to hold the organization's logs, with as few people able to sign in to it as possible and no workloads running in it.
Northgate's security account combines that role with running GuardDuty and Security Hub for everyone, which is common in smaller organizations and acceptable as long as its access is held as tightly as the logs deserve.
The Log Archive Account
Out of reachWhere the bucket lives matters at least as much as how it is configured. The query counts every management event that touches any of Northgate's log buckets, from any account, in the period.
SELECT count(*) AS events_on_log_buckets
FROM cloudtrail_logs
WHERE json_extract_scalar(requestparameters, '$.bucketName')
IN ('northgate-cloudtrail-org', 'northgate-network-logs', 'northgate-dev-cloudtrail')
The query counts management events that name any of the three log buckets. None: nobody changed a log bucket's policy, encryption, lifecycle or versioning in the period, from any account. That is partly design.
The organization trail's bucket and the network logs' bucket sit in the security account, the log archive, where no workload account has access, so the attackers of this course, working in dev and prod, could not reach them even with administrator access in those accounts.
The dev trail's bucket, northgate-dev-cloudtrail, is the exception that proves the rule: it sits in dev, and anyone with administrator access in dev could have changed it.
The count says nothing about the objects inside the buckets, and it is important to be precise about why. The trail does not record data events for its own bucket, both because recording a bucket's object writes into itself creates a loop of events about events, and because the bucket's protection should not depend on the record it protects.
Whether log files were read or deleted is a question for validation, versioning and Object Lock, which the rest of this lesson audits.
The same holds for the other log buckets. northgate-network-logs, which receives flow logs, and the bucket that holds S3's access logs both sit in the security account, so the network and object evidence of Modules 6 and 7 were out of the attackers' reach too.
An audit checks every log bucket's location, not only the trail's, because an investigation is only as protected as the least protected source it relies on.
Location is the archive's first defense, and it is one Northgate got right from the start.
Auditing the Archive
Four configurations read togetherThe bucket's configuration is four separate settings that only mean something together: versioning, Object Lock, the lifecycle rule and the bucket policy. Each is read with its own CLI call, and an audit that reads one or two misses how they interact.
The auditor assembles the bucket's four configurations into one view, in the order an auditor reads them, as a readiness review would gather them from four CLI calls. Read together they tell a story none tells alone. Versioning is on, which keeps the previous version of any overwritten or deleted object for a while, and the CloudTrail write is correctly limited to Northgate's own trail.
But there is no Object Lock, so nothing prevents deletion; noncurrent versions expire after 30 days, so the protection versioning offers is short; after 90 days the files move to a class Athena cannot read without a restore; and one human role can do anything to the bucket, including delete it. Each gap is a single line, and the next sections take them in turn.
The four gaps share a pattern worth noticing. None of them affects the trail's day-to-day work; every log file arrives, every query over recent events runs, and nothing in the console looks wrong. They matter only in the investigations an organization most needs to get right: a late discovery, a suspected tampering, a dispute about what the record shows.
That is why archive settings are so often left at whatever someone chose on the first day.
Each gap also has an order of urgency. The bucket policy is the most urgent, because it is the only one a person can exploit today.
Object Lock comes next, because it protects every file written after it is set and the sooner it is on, the more of the record it covers. Noncurrent versions and the storage class matter on the day of a late investigation, and can follow within weeks.
Kept Long Enough
Retention against needRetention is a question about how far back investigations reach, and it has an honest answer that is usually longer than an organization's first guess. The query measures how far back this course's investigations had to reach.
SELECT min(eventtime) AS oldest_event,
date_diff('day', from_iso8601_timestamp(min(eventtime)), from_iso8601_timestamp(max(eventtime))) AS age_days
FROM cloudtrail_logs
The query reads the oldest event in the record and its age, in days, at the end of the period. Every event in the course is 37 days old or less, comfortably inside the 400 days Northgate keeps its record. The question an audit asks is whether 400 days is enough, and the honest answer depends on how long an attack can go undetected and how long obligations require.
Attackers in long-running intrusions are routinely found many months after they arrived, and an investigation that cannot reach the first access cannot say how the attacker got in or what they did first. Regulations and contracts often set their own minimums. Northgate's 400 days is a reasonable floor; an organization that has never asked the question may find its rule was set to a round number by whoever wrote it.
The same applies to every source, not only the trail. Flow logs, DNS logs, load balancer logs and S3 access logs each have their own bucket and their own lifecycle, and an investigation that can reach back a year in CloudTrail but only a month in flow logs can answer the API question and not the network one.
Retention also needs an exception path. If an investigation or a legal matter needs records past their expiry, the lifecycle rule must not delete them on schedule while the matter is open.
An S3 legal hold on the relevant objects, or a copy into an evidence bucket as 10.6 describes, keeps them; the procedure for applying it should exist before anyone needs it, because a lifecycle rule acts on its own.
An organization that has never had to reach back past 90 days has never felt the restore delay. That is typical of archive settings: they are chosen once, tested by nothing, and found wanting on the one occasion they matter. A short drill, restoring and querying a month of last year's events, is the only way to know how long the real thing would take.
Northgate's choice of 400 days is recorded in its lifecycle rule, so it is easy to read and easy to change; the harder part is agreeing on the number, which belongs to whoever owns the organization's legal and regulatory obligations as well as to the security team.
Readable When Needed
Storage classes and restoresA record that exists but cannot be read on the day it is needed is a delay built into every late investigation. The command below is the validation Module 8 ran, and it meets the same constraint as any query.
aws cloudtrail validate-logs \
--trail-arn arn:aws:cloudtrail:eu-west-2:100000000101:trail/northgate-org-trail \
--start-time 2026-04-14T00:00:00Z --end-time 2026-05-21T23:59:59Z
The command validates the whole period's log files against their signed digests, the step Module 8 took to prove the record unaltered. It reads digests and files from the bucket, so it has the same constraint as any query: files in Glacier Flexible Retrieval have to be restored first. Northgate's lifecycle moves log files to GLACIER after 90 days, which is Flexible Retrieval.
An investigation that needs events from four months ago has to request a restore, typically three to five hours for a standard retrieval, enable restored objects on its Athena table, and only then run its first query. In an incident that is half a working day lost before the first answer, and the first answer is often the one that decides the next step.
Glacier Instant Retrieval is the straightforward fix. It costs more to store than Flexible Retrieval and less than S3 Standard, and its objects are readable at once, with no restore. For a log archive that investigations need to read on demand, that is usually the right class for the whole retention period, with deeper archive classes reserved for records past the point any investigation would plausibly need.
Restores can be planned, and good teams plan them. A team that knows its archive moves to Flexible Retrieval after 90 days can restore a period in bulk at the start of an investigation, at low cost, while the first queries run over recent events.
That works for an investigation that knows its window early. It works badly for one that discovers, partway through, that the attacker was present months earlier, which is exactly when speed matters most.
The cost difference between the classes is real but modest at Northgate's volume, a few hundred megabytes a year.
For an organization recording billions of data events, the calculation changes, and a mix makes sense: management events, which investigations read most, in an instantly readable class for the full retention, and high-volume data events moved to a deeper archive once the period in which they are likely to be needed has passed.
Provably Unaltered
Continuity, digests and Object LockIntegrity has two distinct halves: preventing change, and proving there was none. The query below is the coarsest proof, continuity, and the paragraphs after it take the finer ones.
SELECT count(DISTINCT substr(eventtime, 1, 13)) AS hours_with_events,
date_diff('hour', from_iso8601_timestamp(min(eventtime)),
from_iso8601_timestamp(max(eventtime))) + 1 AS hours_in_period
FROM cloudtrail_logs
The query counts the hours that carry at least one event against all the hours in the period. 884 of 888, and Module 8 explained each of the four quiet hours from other sources.
Continuity is the first, coarse test of integrity, and anyone can run it: deleting a log file usually leaves a hole an hour wide. It cannot detect a file edited in place or replaced with a plausible fake. Digest validation can, because each hour's signed digest lists the hash of every file delivered, and any change to a file, or its removal, fails the check.
Validation proves; it does not prevent, and the difference matters on the worst day. A deleted log file is detected and gone. Object Lock is the prevention: with a default retention in compliance mode, every log file is protected from deletion and overwriting for the retention period, by anyone, including the account's root user.
Together they make the archive's integrity something an investigation can state, not assume. Northgate has validation and not Object Lock, so it can prove whether its record was altered, but not stop someone altering it.
Object Lock has one practical constraint worth planning for in advance. In compliance mode, nobody can shorten a retention period once set or delete a locked object early, including to correct a mistake such as a log file delivered with the wrong content.
That rigidity is the point, and it argues for choosing the retention period carefully before turning the lock on. Governance mode allows a few privileged principals to override it, which is easier to live with and weaker as evidence; for a log archive whose records may be needed in proceedings, compliance mode is the stronger choice.
Object Lock can be turned on for an existing versioned bucket, so Northgate does not need a new archive to fix the gap. Objects already in the bucket are not locked retroactively by a default retention; it applies to new versions from the moment it is set, which is one more reason to set it before an incident rather than during one.
Who Can Change It
The policy that lets a person delete the recordThe last property is about people and their permissions, not settings. A bucket can be versioned, locked and validated, and still be only as trustworthy as the least careful principal its policy allows to change it. The gate below gathers the whole audit into one test.
Expiration Days
Transitions StorageClass
Object Lock retention
Bucket policy actionsRead versioning, Object Lock, lifecycle and policy together.The gate's third clause is the one Northgate fails twice, and both failures are single settings: once for Object Lock, and once for the bucket policy.
The SecurityAdmin role holds every S3 action on the archive. It exists so the security team can manage the bucket, and the convenience is real; but a role that can delete the record is a role an attacker would want, and a role through which an honest mistake can destroy evidence.
The principle for a log archive is simple to state: the trail writes, people read, and nobody deletes, with any administrative access held behind break-glass and alerting, as 0.7 discusses.
Reading access deserves the same care in the other direction. Investigators need to read the archive quickly, from Athena, across all accounts, and a policy that makes reading hard pushes people toward copying logs into places with weaker protection. The goal is a narrow path: a read-only role, used by the investigation workgroup, that can query everything and change nothing, with every use of it recorded by the same trail.
Northgate's SecurityAdmin role is also an example of a broader readiness question that 0.7 takes up: which identities in the organization could destroy the evidence an investigation would need, and how many of them are people. For the log archive the answer should be none in the normal course of work, and a break-glass path, alarmed, for the rare day an administrator must change the bucket itself.
A Procedure for the Archive
Read, retain, readable, changeThe steps below audit any organization's log archive.
aws s3api get-object-lock-configuration --bucket northgate-cloudtrail-orgApplied to Northgate, the four steps find an archive in the right account, versioned, with validation on, kept for 400 days; and four settings to change: Object Lock in compliance mode, Glacier Instant Retrieval for the transition, noncurrent versions kept as long as current ones, and the SecurityAdmin role reduced to read access.
Each is a single configuration change made in the security account, by someone with the authority to make it, and each would matter only on the day an investigation needs the record most. The footer is the step that turns the archive's integrity from a configuration into a statement an investigation can make.
Practice
Three questions about the record.
Next, 0.6 GuardDuty and Security Hub Across an Organization audits how detection is set up for every account.