Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
0.2 How Investigation Works in AWS
Introduction
An investigation on a server starts with the machine: its disk, its memory, its logs. An investigation in AWS starts somewhere else, because in AWS almost everything that matters happens through an API.
Creating a user, stopping an instance, reading an object, changing a policy: each is a request to an AWS service, made by an identity with a set of credentials, from an address, at a time, and AWS records the request. That record is the core of every investigation in this course.
This lesson explains how AWS evidence works: what a single record says and how to read it, who makes the calls in an organization like Northgate, the difference between temporary and long-term credentials, why most of what happens is reading rather than changing, and which services dominate the record. It ends with the questions every AWS investigation starts from.
Nothing here requires writing a query. The queries in this lesson are shown and their results explained, so you can see what the record says before Module 1 teaches you how to ask it. If you have investigated servers before, some of it will feel familiar and some will not, and the differences are where investigators new to AWS most often go wrong.
The figure is the lesson's argument in three steps, and each section below tests one part of it against Northgate's record. Because every change is an API call, the record of calls is the record of what happened.
Because most credentials are temporary sessions, the question is which session made a call and who started it. Because there is no network edge to defend first, identity is where attacks begin and where investigations follow them.
The claim that everything happens through an API needs one important qualification. What happens inside a running instance, a process starting, a file being written to its disk, a command typed in a shell, is not an AWS API call and CloudTrail never sees it.
Those belong to the operating system's own logs and to tools on the instance. This course is about the AWS layer: everything an identity does to AWS resources, which is where cloud attacks most often begin and where the evidence is the same for every organization.
Every Change Is an API Call
One record, read field by fieldThe best way to understand AWS evidence is to read one record slowly, field by field, before reading thousands.
- Who, by kind. A role session, not a user: someone assumed a role to make this call.
- Who, by name. The role is an Identity Center permission set; the session is named for the person.
- With what. ASIA: a temporary key, issued for this session and expiring with it.
- What. The API operation: the action itself, as AWS received it.
- From where. The caller's address as AWS saw it.
- A change. Readonly false: this call changed something.
This is one CloudTrail record, complete, as the Athena table returns it, for an ordinary event: a developer stopping an instance on a Thursday afternoon. It answers the questions an investigator asks of any event. Who: a role session, through Northgate's PowerUserAccess permission set, named for j.eriksen.
With what: a temporary key beginning ASIA, issued for that session 34 seconds earlier. What: StopInstances, on one instance, which moved from running to stopping. From where: an address in Northgate's office range. When: 17:10:44 on 17 April, in eu-west-2. And whether it changed anything: readonly is false.
Some of the fields are worth a second, closer look before moving on. The session was created 34 seconds before the call, which is what a person signing in and going straight to the console looks like. mfaauthenticated is false, because Identity Center handled the authentication before the role was assumed, and the role session itself does not carry it.
Module 3 explains where MFA does appear. useragent reads AWS Internal, as it does for calls made through the console. requestparameters and responseelements are strings holding JSON, which Module 2 teaches you to open.
Every investigation in the course is built from records like this. They are long, and most fields are empty most of the time, but the handful in the margins answer most questions, and Module 1 teaches you to select exactly those.
A record also says, by omission, what it does not know. It does not say why j.eriksen stopped the instance, whether anyone asked her to, or what was running on it. CloudTrail records the request and its outcome; the reasons live in people and in other systems. An investigator reading CloudTrail is reading actions, and has to look elsewhere, often to a person, for intent.
Reading a record this way, field by field, is slow the first few times and quick after that. It is also the habit that prevents the most common early mistake in AWS investigation: reading the event name and the time, and guessing the rest. The identity block alone usually changes what a record means.
Who Makes the Calls
Identities by kindThe identity block's type field says what kind of identity made a call, and across the record the answer is lopsided.
SELECT useridentity.type AS identity_type, count(*) AS events
FROM cloudtrail_logs
GROUP BY 1
ORDER BY 2 DESC
The query counts the record's events by the identity block's type. Role sessions made just over half the calls. Identity Center users account for the sign-in steps, SAML users for the requests that turn a sign-in into role credentials, and AWS services for work they do on Northgate's behalf.
IAM users, the long-term identities with keys or passwords of their own, made 554 calls, and the root user one. Each kind of identity leaves a different trail back to a person.
A role session from Identity Center carries the person's name in its session name, as j.eriksen's did; a session on an instance carries the instance's ID; a session started by an IAM user's key leads back to the user, and from there, only through whoever holds the key, to a person.
A server investigator expects to look for user accounts; an AWS investigator mostly looks for sessions, and for the user, role or key behind each one.
The lopsidedness is a sign of a reasonably well-run AWS organization, and also of how attacks have to work in one.
When people and applications use sessions, an attacker who wants lasting access cannot simply steal a password and wait; they have to steal a session while it is valid, steal the long-term credential that can start new ones, or create one of their own. Each of those leaves a different trace, and Modules 4 and 5 read all three.
Temporary and Long-Term Credentials
ASIA and AKIAEvery call made with an access key carries the key's ID, and its first four letters say what kind of credential it is.
SELECT CASE substr(useridentity.accesskeyid, 1, 4)
WHEN 'ASIA' THEN 'temporary (ASIA)'
WHEN 'AKIA' THEN 'long-term (AKIA)'
ELSE 'none' END AS credential,
count(*) AS events
FROM cloudtrail_logs
GROUP BY 1
ORDER BY 2 DESC
The query sorts every event by the first four letters of its access key. Temporary credentials begin ASIA and are issued by AWS's Security Token Service for a session: when someone signs in through Identity Center, when an instance fetches its role's credentials, when a function runs. They expire on their own. Long-term keys begin AKIA, belong to an IAM user, and work until someone rotates or deletes them.
Northgate's record has 12,041 calls with temporary keys and 390 with long-term ones. The record also shows which IAM users still hold long-term keys: three, for a CI pipeline, a backup job and one developer, each a leftover from before Identity Center. That small number is where much of this course's trouble starts, because a long-term key that leaks keeps working for whoever holds it.
Temporary does not mean harmless, and the course is always careful to say so. A session's credentials can be stolen and used from anywhere until they expire, and Module 7 investigates exactly that: credentials taken from an instance and used from an attacker's server. But the clock is on the defender's side.
A stolen long-term key has no clock at all, and every key of that kind in an organization is worth knowing about before an incident, which is part of what 0.7 and 0.8 assess.
Identity Center's own users are the other half of the identity picture. The 7,848 IdentityCenterUser events are the steps of people signing in, before any role is assumed, and the SAML events are the requests that turn those sign-ins into credentials for a role in an account.
A person's working day appears in the record as a sign-in, a credential request, and then a session's worth of calls under the role.
Every one of those calls is recorded with the same handful of fields the specimen showed, which is what makes them comparable across services.
Reads and Writes
Most of AWS is lookingEvery CloudTrail record says whether its call only read or also changed something.
SELECT readonly, count(*) AS events
FROM cloudtrail_logs
GROUP BY 1
The query counts reads and writes using the readonly field CloudTrail sets on every record. Reads outnumber writes almost three to one. Listing buckets, describing instances, getting policies: most work in AWS starts by looking, and so does most of an attacker's. A stolen key's first calls are almost always reads, to find out what it can reach. Reads have a second, more specific use.
A burst of reads by an identity that rarely reads, especially reads of its own permissions, is one of the most reliable signs that someone new is holding its credentials, and Module 4 reads exactly that pattern as discovery after access. An investigation that filters to writes, because writes look more important, misses the reconnaissance that tells it what the attacker knew and when.
The readonly field is set by AWS, not by the caller, and it follows the operation's nature: Describe, List and Get calls are reads, while Create, Put, Update, Delete, Attach and their kind are writes. A few operations surprise people.
A console sign-in is recorded as a write, because it creates a session; a role assumption is a write for the same reason. Modules 1 and 10 both meet that surprise, in querying and in measuring an attacker's impact.
Writes are the minority, and for exactly that reason they repay reading in full. 6,430 writes in 38 days is fewer than two hundred a day across four accounts, and most of them are the same few operations repeated by the same applications. A write that is not one of those is rare enough to stand out, provided the investigator knows what the usual ones are.
Northgate's record shows the shape clearly because its applications are steady: the portal reads and writes data all day, the backup job writes once a night, the pipeline runs on working days. An organization with busier, more varied workloads has a noisier record, and the same habit, knowing the usual before judging the unusual, matters even more.
The Services in the Record
Identity and data dominateThe services that appear most in the record are the ones every investigation passes through.
SELECT eventsource, count(*) AS events, count(DISTINCT eventname) AS operations
FROM cloudtrail_logs
GROUP BY 1
ORDER BY 2 DESC
The query counts events and distinct operations by the service that recorded them. Sign-in and STS together are identity in motion: people signing in, roles being assumed, credentials being issued. S3 and KMS are data: objects read and written, and the encryption keys that protect them. EC2, Lambda and IAM follow. Fourteen services in all, sixty-one different operations, across four accounts.
Seven of the fourteen services are ones a typical investigation reads every time: sign-in, STS and SSO for identity, IAM for permissions, S3 and KMS for data, and EC2 for compute. The others appear when an attack reaches them. Each service records its calls in its own way, with its own fields in the request and response, and much of Modules 1 to 3 is learning to read them.
Some services in the record are there only because of others. KMS appears 3,496 times, every one of them S3 asking KMS to encrypt or decrypt an object for an application or a backup job; the application called S3, and S3 called KMS on its behalf.
Reading those calls as separate actions would double-count the work; reading them as one request and its consequence is how an investigator keeps the record in proportion, and Module 6 uses exactly that link as a witness to what data was read.
Fourteen services is a small number for a real organization of Northgate's size, which might show fifty or more, and that is one of the ways the sample is kept readable. The method does not change with the number: find the services that touch identity, data and compute, read those first, and add the others as the investigation reaches them.
The Question and Its Record
Is it an API call?Most investigative questions in AWS can be turned into a question about API calls, and the habit of making that translation first saves more time than any query technique, and the gate below makes that the first step.
eventname = 'StopInstances'Name the API call; find it in CloudTrail.The example in the gate is the specimen of Section 1: a question about who stopped an instance, answered by one record. When the answer is yes, CloudTrail holds it, provided the trail recorded that kind of call.
When it is no, another source has to: a network connection is in the flow logs, a name lookup in the resolver logs, a request to a web application in the load balancer's logs, a read of an object in S3's access logs if data events were not recorded.
What an organization turns on decides which of those questions it can ever answer, which is why the rest of this module audits exactly that.
The gate also exposes a dependency that matters long before any incident. Each no is a source the organization must have turned on, kept and made queryable in advance; none can be switched on after the fact to cover the past. That is the case for the readiness work of this module, and the reason Module 0 comes first.
The no branch is more common than people new to AWS expect at first. A request to the customer portal that made it reach out to the internet, a DNS lookup of a strange name, a read of an object in a bucket whose data events were never selected: each is a question an investigation in this course asks, and none of them is answered by CloudTrail's management events.
Knowing which source answers which question, before the question arises, is most of what investigation readiness means.
What Changes From a Server
Unlearning, and keepingInvestigators who come to AWS from servers bring habits worth keeping and some worth setting aside, and naming them early saves a lot of confusion later.
From a server to AWS
What changes for an investigatorThe record above sets the two side by side. The discipline is the same, and it is the part that transfers: a timeline, a scope, evidence for every claim. What changes is where the evidence lives, who keeps it, and what it describes.
There is no disk to image for an API call; there is a record of the request, kept by AWS, for as long as the organization chose to keep it. And there is no host to isolate for a stolen key; there is a credential to stop, which Module 10 makes a lesson of its own.
Two things carry over from server work unchanged. The first is the timeline: an AWS investigation still lines up events in order and asks what happened between them, and Module 2 builds one from the record. The second is skepticism about what the evidence says.
A source address can be a proxy, a user agent can be set to anything, a session name is whatever the caller chose; CloudTrail records them faithfully, and an investigator still has to decide what each is worth.
What changes most is speed, in both directions. A terminated instance or a deleted function is gone in seconds, and nothing on AWS's side keeps a copy for the customer. An attacker can create resources in a Region nobody watches and remove them before anyone looks.
The record of the API calls survives, if the trail kept it; the resources themselves often do not, which is why Module 10's lesson on preservation puts copying ahead of almost everything else.
None of this makes servers or their evidence irrelevant. Many AWS incidents reach an instance, and once they do, the operating system's evidence matters as much as it ever did. The course's focus on the AWS layer reflects where cloud attacks begin and where investigators new to AWS have the most to learn, not a claim that the layer below does not count.
The First Questions
A procedure for any AWS investigationThe steps below are the first questions of any investigation in AWS, whatever its subject.
eventname IN ('StopInstances', ...)useridentity.arn, useridentity.accesskeyidThe procedure is short because the questions are general and few; the work is in answering them for a particular incident, which every later module does in detail. It is also the procedure to use when someone hands you an alert in an account you have never seen: name the operations, find the session, check the credentials, then ask what else is recorded, before forming any view of what happened.
Applied to the record, the four questions are how every investigative module of this course begins: which operations, which session, which credentials, and which other sources hold the rest. The footer is the limit every investigation meets, and it is the reason this module turns next to what an organization should record before an incident.
Practice
Three questions about the record. Each needs a count; Module 1 explains every part of the queries you will write.
Next, 0.3 Organization Trails audits the trail that records every Northgate account, and the gaps in it.