Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
0.4 Data and Network Activity Events
Introduction
An organization trail records every account and every Region, and still leaves most of what happens to data unrecorded. By default, and at no CloudTrail event charge for the first copy, a trail records management events: the calls that create, configure and delete resources.
The calls that read and write the data inside those resources, objects in buckets, invocations of functions, items in tables, are data events, and a trail records them only for the resources its selectors name.
Calls through VPC endpoints are a third kind, network activity events, recorded only for the services selected. Both cost money per event, so organizations choose, usually once, when the trail is set up, and every choice not to record is permanent: data events cannot be recovered for a period in which they were not selected. The trade between cost and coverage is real.
A busy bucket can produce millions of data events a day, and recording every object operation in every bucket in a large organization costs more than many security budgets allow. But the cost of not recording falls entirely on the day of an incident, when the question is which objects left and the answer is that nobody kept the record.
Choosing well means knowing which resources hold data or run code an investigation would need to account for, and recording those. This lesson measures what Northgate's choices recorded and what they left out, audits the trail's selectors, and sets out how to decide resource by resource.
The figure is the three kinds of event, and the band beneath them is the one fact that makes the choice urgent. Management events tell an investigation how things were set up; data and network activity events tell it what was done with them. Northgate records the first everywhere, and the other two only where it chose.
The three kinds of event share one record and one table in Athena, distinguished by the eventcategory field, and Module 1 shows how to separate them in a query. They differ in what triggers them, what they cost, and what they can answer, and the audit in this lesson is about the last of those.
Every number in this lesson describes Northgate's choices as they stood throughout the period; none changed between 14 April and 21 May, which Module 8 confirmed by finding no change to the trail's selectors in the record.
Data Events in the Record
A minority, and the part that mattersData events are a small part of the record, and in a data theft they are the evidence. The first question is how much of the record they make up.
SELECT eventcategory, count(*) AS events
FROM cloudtrail_logs
GROUP BY 1
The query counts the record's events by category, management or data. Most of Northgate's 23,570 CloudTrail events are management events, and every account records them because the trail does so by default. The 3,549 data events exist only because someone chose the buckets they come from.
In the investigations of this course, the management events say who changed a bucket's policy; the data events say who read the backups after it changed. The second question is usually the one that matters to the people whose data it was.
Data events carry a lot per record. Each says which object, by key, was read or written; the identity that did it; the address it came from; whether the request succeeded; and, for reads, how many bytes left.
That is what makes them the evidence of a data theft. A management event tells an investigator that a bucket's policy was opened to the public. Only data events, or S3's own access logs, can tell them which objects anyone then read.
Data events also record failures, which is easy to overlook. A read refused because the caller lacked permission is recorded with its error, which turns the data events of a protected bucket into a record of who tried to read it as well as who did. For a bucket an attacker is probing, the refusals can be the first evidence of interest, long before anything succeeds.
Three Buckets Recorded
What the selectors choseNorthgate selected data events for three of its six buckets. The query below shows which, and how much each recorded over the period.
SELECT json_extract_scalar(requestparameters, '$.bucketName') AS bucket, count(*) AS data_events,
count(DISTINCT eventname) AS operations
FROM cloudtrail_logs
WHERE eventcategory = 'Data'
GROUP BY 1
ORDER BY 2 DESC
The query counts data events and distinct operations by bucket. Three buckets, three different purposes: the portal's data, the database backups, and the CI pipeline's build artifacts. Each was chosen for a reason. The portal's data is what customers entrust to Northgate. The backups are the organization's ability to recover.
The artifacts are what the build pipeline produces and deploys, and an attacker who could swap one would compromise every system that installs it. Those three choices are why Module 4 can see what the leaked CI key normally uploads, and why Module 6 can count exactly which backups were read on 16 May.
The three are named individually, by ARN prefix. That has a consequence an audit has to remember: a bucket created tomorrow is not recorded until someone adds it to the list. An organization that creates buckets often may prefer a selector that covers every bucket in an account, or every bucket except a few named exclusions, which costs more and protects against the bucket nobody remembered.
Recording the backups bucket was the choice that mattered most in this course's period. On 16 May an attacker with the backup user's key read every backup in it, and the record of those reads, each object and each byte count, is how Module 6 established what was taken.
Had the bucket been left out, the investigation would have known the policy was opened and the encryption changed, from management events, and would have relied on S3's own access logs, best-effort and less precise about callers, to establish what left.
The three choices were the right ones for the period's attacks; the next section shows they were not enough for every attack an organization should expect.
Two Buckets Not Recorded
Bucket-level calls onlyTwo of Northgate's buckets have no data events, and the record shows exactly what that means for an investigation: the buckets appear, and their contents do not.
SELECT json_extract_scalar(requestparameters, '$.bucketName') AS bucket, count(*) AS management_events,
array_agg(DISTINCT eventname) AS operations
FROM cloudtrail_logs
WHERE eventsource = 's3.amazonaws.com'
AND eventcategory = 'Management'
AND json_extract_scalar(requestparameters, '$.bucketName') IN ('northgate-dev-scratch', 'northgate-app-assets')
GROUP BY 1
The query lists the management events for the two buckets without data events, by operation. Both buckets appear, but only through bucket-level calls: the audit function reading each bucket's policy and public access settings every day.
Not one object operation appears, because none was recorded. northgate-app-assets holds the public website's images and scripts, which anyone on the internet can read, and leaving it out is a reasonable decision: recording the public's reads would cost a great deal and show little. northgate-dev-scratch is different.
It is a scratch space in the dev account, used for test exports among other things, and data put there is exactly the kind an attacker would copy out. That omission is a gap.
S3 server access logs are a partial alternative. A bucket can deliver its own access logs, a record of every request to it, to another bucket, without CloudTrail's per-event charge.
Northgate turns them on for some buckets, and Module 6 reads them beside the data events. They are delivered on a best-effort basis and identify callers less precisely than CloudTrail does, so they complement data events rather than replace them, but for a bucket like dev-scratch they would be far better than nothing.
The dev-scratch gap is also a reminder that dev accounts hold data. Test exports, copies of production tables used for debugging, configuration files with credentials in them: all routinely end up in scratch buckets in development accounts, and all are as valuable to an attacker as the production originals. A data event policy that records production and ignores development records only half of where the data is.
Functions That Run Unrecorded
Lambda invocationsS3 is not the only service with data events that matter to an investigation. Lambda functions run code with the permissions of their roles, and whether a function ran, and who made it run, is a data event too.
SELECT eventname, count(*) AS events
FROM cloudtrail_logs
WHERE eventsource = 'lambda.amazonaws.com'
GROUP BY 1
ORDER BY 2 DESC
The query lists every Lambda event in the whole record by operation. Northgate's functions run every day, but nothing in the record says so: function invocations are data events, and the trail selects none for Lambda. The functions appear only when someone lists or describes them, and one appears when it was created.
That matters for the staged function of 18 May. Module 10 had to prove it never ran by a roundabout route, by showing that Lambda never assumed the role it would run as. With Lambda data events, every invocation, and the identity that triggered it, would be in the record directly.
Lambda's data events have the same trade-off as S3's. A function called thousands of times a minute by an application generates a matching number of events, and an organization might record only production functions, or only functions whose roles can reach sensitive data. Northgate runs only a handful of production functions, so recording them would cost little. The decision not to was simply never made.
There is a broader point in the staged function. Every resource an attacker can create is a resource whose use the record may not show. A user's creation is a management event; its use is recorded too, because its calls are management events.
A function's creation is a management event; its use is a data event. Knowing which kinds of resource fall on which side of that line is what lets an investigator say, during an incident, whether absence of evidence means the thing did not happen or only that nobody was recording.
Other services have data events too, and an audit should know which of an organization's resources fall among them: DynamoDB items, for example, and a growing list of other resource types recorded through advanced selectors.
Northgate uses none of those services in this period, which is why the course's audit stops at S3 and Lambda. An organization that stores its secrets or its customer records in another service has to ask the same question of each.
Network Activity Events
Who used an endpointNetwork activity events are newer, generally available since February 2025, and answer a question nothing else does: who used a VPC endpoint, and who it turned away.
SELECT eventsource, vpcendpointid, count(*) AS events,
count_if(errorcode = 'VpceAccessDenied') AS refused_by_endpoint_policy
FROM cloudtrail_network_activity
GROUP BY 1, 2
The query counts network activity by service and endpoint, and how many calls the endpoint's policy refused. A VPC endpoint lets instances reach an AWS service without going through the internet, and an endpoint policy can restrict which principals or resources may be used through it.
Network activity events record each call through the endpoint, from the endpoint's side, including calls made with credentials from another organization and calls the policy refused. The 18 refusals in Northgate's record are calls the policy stopped, evidence of attempted use that management and data events would never show. The prod VPC also has a KMS endpoint, and nothing records its use.
Network activity events also close a gap data events leave. Data events record an object read from the bucket's side, whoever the caller was; network activity events record the call from the network's side, through the endpoint, with the endpoint policy's decision.
Together they can show a caller inside the VPC using credentials from outside the organization, the pattern of stolen credentials being used from a compromised instance, which neither source shows alone.
The KMS endpoint matters for the same reason the S3 one does. Every read of an encrypted object in Northgate's prod buckets involves KMS, and an instance in the prod VPC reaches KMS through that endpoint.
Recording its network activity would show every decryption request routed through the VPC, and any refused by the endpoint's policy. The omission is cheap to fix, and nothing in the record or the configuration says why it was made, which is itself a sign that it was not a decision.
Network activity has one more property worth knowing: it can be set to record only refused calls. That keeps the cost low for a busy endpoint while still capturing the evidence that matters most, the attempts the endpoint's policy stopped. For Northgate's KMS endpoint, a refused-only selector would be a cheap start.
Auditing the Selectors
What the trail actually recordsThe trail's selectors are where every one of these choices is written down, in one place, and auditing them is how an organization finds out what it actually decided.
The auditor shows get-event-selectors exactly as the CLI prints it for the organization trail. Northgate uses advanced event selectors, which are required for network activity events and let each selector filter on fields such as the event category, the resource type and the resource's ARN.
Three gaps follow from what the queries showed: a bucket holding data with no data events, a resource type, Lambda, with none at all, and an endpoint service left out of network activity.
Two lines look like candidates and are not: the management selector, which is exactly right, and the artifacts bucket, whose data events Module 4 depends on. And one bucket is missing by design, the public assets bucket, which is a decision rather than a gap.
The audit's findings are modest in number and large in consequence.
None of the three gaps cost Northgate an investigation in this course's period, because no attack happened to pass through the dev scratch bucket, a function invocation or the KMS endpoint. Each is nonetheless a route an attacker could use with no data trail, and the course's attackers showed they would use whatever route was open. Readiness is about the next incident, not the last one.
The decoys deserve as much attention as the gaps, perhaps more. An auditor who flags the management selector because it has no exclusions, or the artifacts bucket because builds seem unimportant, would weaken the record while trying to tidy it. Every line in a selector is a decision someone made, and the audit's job is to find the decisions that were wrong, not to second-guess the ones that were right.
Fixing all three is a single update to the selectors, made from the management account: the scratch bucket's ARN added to the list, a second data selector for Lambda functions, and the KMS service added to network activity. The update is itself a management event, PutAdvancedEventSelectors, recorded with the new selectors in its request, so the record will show when the coverage changed and who changed it.
After the update, the next readiness audit rereads the selectors and checks the three changes held, the same way it checks the trail's settings.
Deciding Each Resource
A question to ask of everythingThe audit becomes a decision for each resource, made explicitly, and the gate below is the question that decides it.
resources.ARN StartsWith
resources.type AWS::Lambda::Function
eventSource kms.amazonaws.comAsk the question of each resource, and check its selector.The gate's question is about consequences for an investigation, not categories of resource. A bucket of public images fails it; a scratch bucket where anyone might put an export passes it.
A function that reformats logs fails it; a function that runs with a role able to read production data passes it. The answer is not always yes, and the gate asks for the no to be written down, because the most damaging gap in a record is one everyone assumed was covered.
Cost is the usual reason for a no, often unexamined, and it is worth estimating rather than assuming. The management events in a bucket's account show roughly how busy it is; a quick count of its objects and their change rate, or a few days of S3 access logs, gives the order of magnitude of its data events.
For many buckets the answer is smaller than feared, and for the few where it is large, a selector that records only writes, or only one prefix, can keep the evidence that matters at a fraction of the cost.
The record of decisions belongs with the selectors. A short document listing every bucket, function and endpoint, recorded or not, with a reason for each no, is what lets the next audit check whether anything changed, and what lets an investigator during an incident say with confidence that a resource was never recorded rather than wonder whether it was missed.
Whichever way the decision goes, the selectors should be reread whenever the organization changes shape: a new account, a new kind of workload, a new VPC endpoint. Each brings resources the existing selectors may not cover.
A Procedure for Selectors
Inventory, read, decide, costThe steps below choose data and network activity events for any organization.
aws s3api list-bucketsaws cloudtrail get-event-selectors --trail-name northgate-org-trailApplied to Northgate, the four steps find six buckets, three recorded and one deliberately excluded, leaving northgate-dev-scratch to add; Lambda functions with no invocation record, to add for production functions at least; and a KMS endpoint to add to network activity. Each addition costs per event, and each answers a question the course shows an investigation asking.
The footer is the choice between naming resources and covering them by prefix or account, which decides whether the next bucket is recorded automatically or only when someone remembers.
Practice
Three questions about what the trail recorded.
Next, 0.5 Retention, Integrity and the Log Archive audits how long the record is kept and whether it can be trusted.