In this section

0.3 Organization Trails

Module 0

Introduction

CloudTrail records management events in every account by default, for ninety days, in each account's Event history. That default is not enough for an investigation. It is kept for too short a time, it lives inside each account where anyone who controls the account can see it, and it has to be read account by account and Region by Region.

A trail fixes all three: it delivers every event to an S3 bucket, where it can be kept as long as the organization chooses and queried with Athena. An organization trail goes further. Created once in the management account, it records every account in the organization, including accounts added later, in every Region, into one bucket, and no member account can stop or change it.

The decisions a trail embodies are made once, usually by whoever set up the organization, and then forgotten. They are also the decisions every later investigation depends on: whether the record covers the Region an attacker used, whether it includes the IAM change that granted them access, whether it can be proven unaltered.

An investigator who knows them before an incident knows what the record can answer; one who learns them during an incident learns them from the questions it cannot.

This lesson audits Northgate's trails: the organization trail the whole course reads, and an older trail in the dev account that nobody removed. It shows from the record why each of the trail's settings matters, and it ends with the five settings any organization's trail should have.

An organization trail Management account creates and owns the organization trail Management every Region Security every Region Prod every Region Dev every Region northgate-cloudtrail-org in the security account One trail, created where member accounts cannot change it, recording every account in every Region into a bucket none of them controls.

The figure is the organization trail's shape, as Northgate built it. It is created in the management account, it records every account in every Region, and it delivers to a bucket in the security account. Each part of that shape answers an attack the course investigates.

A note on cost, since it shapes these decisions in practice more than anything else. The first copy of management events in each Region is delivered by a trail without CloudTrail charging for the events themselves; the organization pays for the S3 storage and for any queries.

Data events and network activity events are charged per event, which is why they are chosen resource by resource, and why 0.4 is a lesson about choosing well.

01

What an Organization Trail Is

One trail, every account

Northgate's organization trail is the record behind every query in the course.

northgate-org-trail

Management account, eu-west-2

Scope

Organization trail: every account, present and future

Regions

Every Region, with global service events

Integrity

Log file validation on: hourly signed digests

Delivery

northgate-cloudtrail-org, in the security account

Insights

Off

The record above is the trail's configuration in brief, as describe-trails and the trail's details page report it. It was created in the management account and marked as an organization trail, so AWS applies it to every member account automatically. It is multi-Region and includes global service events, so it records everywhere. It validates its log files, so they can be proven unaltered.

It delivers to a bucket in the security account, so no member account controls its record. Insights, which flags unusual rates of API calls or errors in management events, is off; that is a choice worth revisiting, not a gap that affects what the trail records. Insights is a detection feature, and the course builds its detections from the record instead, in Module 9.

Event history still has a genuine use alongside the trail. It is the quickest view of the last ninety days of an account's management events, needs no setup, and is where an investigator with no Athena access can start. But it is per account and per Region, it keeps nothing older, and anyone who controls the account controls who can see it.

A trail is what makes the record the organization's rather than each account's.

Each of those choices was available to Northgate from the start, and each is a single setting. None of them is hard to make; all of them are easy to leave unmade, because a trail without them still records something and looks as if it works.

02

Every Region

Including the ones nobody uses

The multi-Region setting matters most, perhaps surprisingly, for the Regions an organization does not use.

SELECT awsregion, count(*) AS events
FROM cloudtrail_logs
GROUP BY 1
ORDER BY 2 DESC

The query counts the record's events by Region. Almost all of Northgate's activity is in eu-west-2, which makes it tempting to record only there. 542 events say otherwise. 517 are in us-east-1, where CloudTrail records global services such as IAM: every user created, every key issued, every policy changed.

The rest are in Regions Northgate does not use at all, 21 in eu-north-1 and one each in four others. Those few events are exactly the ones an investigation most needs, because an attacker who works where nobody looks is counting on a trail that does not look there either.

Recording every Region costs little in practice, because a Region with no activity produces no events to pay for. The cost of not recording it is paid only during an incident, and it is paid in full: an attacker who launched instances or created keys in a Region the trail ignored has left no record that anyone can recover.

AWS can restrict which Regions an organization uses with a service control policy, which is a separate decision; the trail should record every Region either way, so that a call refused by such a policy is still on the record.

The Regions with a single event each deserve a word, because they show why counts alone mislead. One event in a Region looks like noise.

In Northgate's record, four of those single events came from the same key within seconds of each other, in four Regions it had never used, and Module 8 reads them as an attacker checking where else it could work. A trail that recorded them is what makes that reading possible at all.

03

Every Account

Including the ones added later

The organization setting matters both for the accounts an organization has and for the ones it will add.

SELECT recipientaccountid AS account, count(DISTINCT awsregion) AS regions, count(*) AS events
FROM cloudtrail_logs
GROUP BY 1
ORDER BY 1

The query counts each account's events and the Regions they occurred in. One trail records all four accounts. Dev shows activity in seven Regions, more than any other account, which is the kind of fact an investigator wants to see across every account at once rather than discover account by account.

When Northgate adds an account, for a new product or an acquisition, the organization trail records it from the moment it joins. A trail per account would rely on someone remembering to create each one, correctly, every time.

Every account means the management account too, which is easy to forget because it feels like infrastructure rather than workload. It is where Identity Center lives, where every person's sign-in begins, and where the organization's own policies are changed, and the organization trail records it like any other. A trail created only in member accounts would leave out the account that controls all the others.

The account view is also where gaps in coverage would usually show first. An account with no events for a day it was in use, or a Region with events in one account and none in another that uses it, is worth a question before it is an incident. Northgate's record has no such gap, which Module 8 checks hour by hour.

04

Out of Reach

Member accounts cannot change it

The most important property of an organization trail is one the record demonstrates directly, under attack.

SELECT eventtime, eventname, recipientaccountid AS account, errorcode
FROM cloudtrail_logs
WHERE eventsource = 'cloudtrail.amazonaws.com'
  AND readonly = 'false'

The query lists every call that tried to change anything in CloudTrail itself. Two calls from the dev account, two and a half minutes apart, tried to change the organization trail: one to stop it logging, one to delete it. Both were refused.

A member account can see an organization trail, but only the management account, or an account the management account delegates to, can stop, change or delete it. The identity that made those calls had full administrator access in dev, and it still could not touch the trail. Module 8 investigates the attempts; for readiness, they are the evidence that the design works.

The protection has a boundary worth knowing precisely. It covers the trail's configuration, not the log files once delivered: those sit in the trail's bucket, and whoever can write to that bucket can delete or overwrite them.

That is why the bucket's location matters as much as the trail's, and why Northgate's is in the security account rather than the management account or any workload account. Who can reach that bucket, and whether it can prevent deletion, is 0.5's audit.

The other boundary is the management account itself. Whoever controls it can stop the organization trail, as AWS designed, because someone has to be able to manage it. Access to the management account is therefore part of the trail's protection, and the reason 0.7 audits who holds it.

Delegation is worth considering for the same reason, though it is a trade-off. CloudTrail lets the management account register a delegated administrator, a member account that can manage organization trails.

Northgate has not, so only the management account can change its trail; an organization that delegates to its security account keeps day-to-day trail management away from the account that controls everything else, at the cost of one more account that can stop the record.

Delegating or not, the management account remains the trail's ultimate owner, and the people who can sign in to it are, in effect, the people who can turn the organization's record off.

05

A Trail Left Behind

Auditing what dev can see

Organizations that adopted organization trails after years of AWS use often still have the trails they started with, one per account, each set up by whoever ran that account at the time.

The auditor shows describe-trails exactly as it would print in the dev account. Mark the faults you would raise, then compare with the explanation. Run in the dev account, describe-trails lists two trails: the organization trail and dev-account-trail, created before the organization trail existed and never removed.

Every setting that makes the organization trail trustworthy is weaker in the dev trail. It records one Region and no global service events, so it would miss both the escalation's IAM changes and the activity in other Regions. It has no validation, so its files could be altered undetectably. And it writes to a bucket in the dev account, under the control of whoever controls dev.

Leftover trails like this one are common, for a simple reason: removing one feels risky, and keeping it seems harmless. Somebody might be reading its bucket; an old script might depend on it; it costs little. The real danger is not that the dev trail exists; it records nothing the organization trail does not.

The danger is that someone, during an incident, reads it instead: it is the trail in the dev account's own console, its bucket is the one dev's engineers know, and it looks like the account's record. An investigation built on it would be missing the most important events and could not prove the rest unaltered.

A member trail is not always wrong, and the audit should not assume it is. An account owned by a separate business unit with its own security team, or an account whose logs must by contract be kept somewhere specific, may have a legitimate reason for its own trail.

The test is whether the trail has a named owner and a reason that survives the question of why the organization trail is not enough. Northgate's dev trail has neither; it predates the organization trail and nobody removed it.

Keeping it also has a running cost that is very easy to miss: the dev trail is a second copy of the dev account's management events in eu-west-2, and CloudTrail charges for every copy after the first. Northgate pays, every month, for a weaker duplicate of a record it already has. Removing it has a small, real cost of its own.

The trail's bucket holds the dev account's events from before the organization trail began, and those might matter to a question about the past. So the bucket is kept, read-only, for as long as its contents might be needed, and the trail is deleted, which stops new writes and removes the second, weaker record of the present.

Each of the four gaps would have cost a specific investigation in this course. Without other Regions, the CI key's sweep of four Regions on 11 May would be invisible. Without global service events, the escalation's new user, key and administrator policy on 18 May would not exist in the record.

Without validation, the statement Module 8 makes, that the record of the evasion attempts is complete and unaltered, could not be made. And with the bucket in dev, the attacker who held administrator access in dev on 18 May could have deleted the account's own record of what they did.

06

The Five Settings

A test for any trail

The audit reduces to a test any organization can apply to the trail it relies on, from the CLI or the console, in a few minutes.

A trailAny trail an organization relies onnorthgate-org-trailAll accounts
The gateIs it an organization trail, multi-Region, with global service events and validation, delivering to a bucket no member account controls?
"IsOrganizationTrail": true
"IsMultiRegionTrail": true
"IncludeGlobalServiceEvents": true
"LogFileValidationEnabled": true
Read describe-trails from the management account and from a member.
All fiveRely on it; audit what it records next (0.4).
Any missingFix that setting, or replace the trail with one that has all five; remove member trails that duplicate it weakly.
An investigation trusts one trail. That trail has to cover everything and be out of reach of what it records.

The five are the settings from the figure, plus the bucket's location, and each has a field in describe-trails. Northgate's organization trail passes all five. The dev trail fails four. The fix for the dev trail is removal, not repair, after checking that nothing reads its bucket, because an organization trail that passes the test makes member trails redundant, and a redundant weaker record is a liability.

Passing the test settles where the record is, whether it is complete in scope, and whether it can be trusted. It does not settle what the record contains, which depends on choices made inside the trail's scope, and those are where Northgate's real gaps lie.

The test also applies to the organization trail's future, which nobody audits by default.

A setting changed by mistake, validation turned off during a cost review, the trail narrowed to one Region to save money, would weaken it without anyone noticing. Each such change is itself a management event in the management account, UpdateTrail with the new setting in its request, and a detection on it, of the kind Module 9 builds, turns the audit from a one-off into a standing control.

Where the dev trail is concerned, the decision is now an action with an owner: the dev account's lead confirms nothing reads northgate-dev-cloudtrail, the trail is deleted, and the bucket is set to read-only and kept for as long as its older contents might be needed. Module 10's review is where actions like this one are tracked to completion.

07

Is It Logging?

Configuration is not operation

The last check in the procedure is the one most often skipped, because configuration feels like proof. A trail with the right settings can still be stopped, misdelivering or failing, and its configuration would not show it.

aws cloudtrail get-trail-status --name arn:aws:cloudtrail:eu-west-2:100000000101:trail/northgate-org-trail

get-trail-status reports whether the trail is logging, when it last delivered a log file and a digest, and any delivery errors, such as a bucket policy that refuses the trail's writes.

A readiness audit reads it as well as the configuration, because a trail that stopped delivering a month ago looks exactly like a healthy one in describe-trails. For Northgate, Module 8's continuity check showed events in every hour of the period, which is the record's own evidence that the trail was logging throughout.

Status also belongs on a schedule, not only in an audit. A readiness audit done once tells you the trail worked on the day of the audit; a check of get-trail-status, run daily and alerting on any delivery error or any change to IsLogging, tells you it has worked every day since.

The stop and delete attempts of 18 May were refused, but an organization should want to know the moment anyone tries, which Module 9's tampering detection does from the record itself.

In the CloudTrail console the same facts are on the trail's details page: the trail's status at the top, the last delivery times and any error under its logging section. In an organization with many accounts the CLI is quicker, and the same command run against each account is how an audit finds a member trail that stopped long ago.

For an organization without Module 8's continuity check, the simplest equivalent is a daily count of events per account from the trail: an account that suddenly records none, while its applications are known to be running, is a trail problem until shown otherwise.

08

A Procedure for Trails

List, check, confirm, decide

The steps below audit the trails of any AWS organization, from the smallest to one with hundreds of accounts; only the number of describe-trails runs changes.

Auditing an organization's trailsCloudTrail console and CLI
01List every trail, from every account
describe-trails in the management account and in each member.
Most often skipped
aws cloudtrail describe-trails
02Check the organization trail's five settings
Organization, multi-Region, global events, validation, bucket location.
03Check it is logging
get-trail-status: IsLogging true, recent delivery, no errors.
aws cloudtrail get-trail-status --name northgate-org-trail
04Decide each member trail's fate
Remove duplicates that are weaker; keep only with a reason.
Still to do: what the trail records inside its scope. A trail can cover every account and Region and still leave out the events an investigation needs, because data and network activity events are chosen one resource at a time; 0.4 audits that.

The audit's output is short: a list of trails, each with its five settings and its status, and a decision for each one that is not the organization trail.

Applied to Northgate's four accounts, the four steps find one organization trail that passes every check, confirmed logging by the record's continuity, and one member trail in dev that fails four of five and should be removed. The footer is the next question: what the organization trail records inside its scope, which 0.4 audits.

Practice

Three questions about what the trail recorded.

Next, 0.4 Data and Network Activity Events audits what the organization trail records inside its scope: which buckets, functions and endpoints, and which it leaves out.