In this section

0.8 What This Course Builds

Module 0

Introduction

This course builds one skill: turning a security question into an SPL search, and the search into an answer that can be defended. Ten modules build it, each adding a set of commands and the habits to use them well, and every module is taught on the same month of events. The month holds several attacks, and finding them is how the course tests what it teaches.

This section sets out what each module adds and previews the attacks. You will see an account worn down by repeated prompts, a service account taking files from SharePoint in eighteen minutes, a laptop's recovery options removed in three, 432 MB leaving through the firewall, a beacon's single large upload, and the one guess that worked.

By the end you should know what the course will ask of you, what each module adds, and what you will be able to do at the end of it.

The attacks the course finds in the month 2 Mar guessing r.scott 5 Mar VPN and cloud spray 19 accounts 8 Mar upload beacon 268 MB 12 Mar 11:02 ransomware prep NE-BENNETT-LT 12 Mar 16:03 SharePoint downloads svc-netops 12-14 Mar transfer out 432 MB Ten modules of SPL, one month of events, and these attacks to find in them.

The timeline runs from 2 March to 14 March, twelve days in which the month's quiet routine is interrupted again and again. The prompts against d.foster, on 1 March, come the day before it begins, and Section 1 shows them. The first half of the month is mostly ordinary activity, which is what makes the second half stand out.

The attacks were built into the month so that every module has something real to find, and the course does not reveal them all at once. Each is introduced where the commands that find it are taught.

The diagram is the month's attacks in time. None of them is hidden by design; each is in the events, and each is found by searches a careful analyst would write.

What each module builds

The course, Modules 1 to 10

1 Search Structure and Fields

questions as searches; fields in any sourcetype

2 Searching and Filtering

terms, Boolean logic, time ranges, addresses

3 Shaping Results

eval, labels, time formats, tables

4 Statistics

stats, rates, distributions, timechart

5 Combining Data

lookups, subsearches, joins, correlation

6 Text and Multivalue Data

rex, JSON, multivalue fields, decoding

7 Macros, Lookups and Knowledge Objects

reusable searches and definitions

8 Time and Sequence

bins, baselines, windows, sessions

9 Data Models and tstats

the CIM, tstats, acceleration, coverage

10 Search Optimization and Validation

cost, limits, checking any search, SPL2

Modules 1 to 4 are the foundation, and the course's pace in them is deliberate: every later module assumes their commands without explaining them again.

The module names are the course's table of contents, and each module's summary says what it built, so the card can be read again at any point to see where the course is.

The modules follow the ladder from Section 0.7, one module per rung, with reading and filtering split across the first two. Each module's searches use the commands of the modules before it, so the order matters more in this course than in many.

The card lists the ten modules. Each has eight sections, nine in Module 10, and a summary with a knowledge check and a set of questions no section answered, so that each module ends by testing what it built.

01

The First Four Modules

From question to answer

Modules 1 to 4 build the core: a question as a search, the fields of any sourcetype, filtering, shaping and statistics. Most searches an analyst writes every day use only these.

index=azuread category=SignInLogs src_ip=91.215.85.201
| stats count by user action

The success came at the end of the run of failures, not mixed among them. Sorting such events by time, Module 3's sort and Module 8's sequences, shows the order in which the user gave way.

One success after fourteen failures, at 14:12 on 1 March, is the moment the attack worked. From then on the attacker held a session for d.foster, and every later event from that account is a question for the investigation.

The pattern is the opposite of the spray: one account, many attempts, one address, and an ending in success. A count by account and outcome separates the shapes of attack at a glance.

The address's 16 events are all for d.foster: 14 failed sign-ins, one successful one, and one more event in another Entra ID category. A complete count by action, without a filter, shows everything an address did.

91.215.85.201 is the address that sent the prompts. In the identity data, it appears only for d.foster, which is part of what makes it suspicious: an address that tries one account repeatedly and nothing else.

Fourteen failures and one success for one account from one address. A count by account and outcome, Module 4's first tool, is enough to see it.

Counting by action, rather than filtering on it, is the cheapest way to see everything an address did. It is also the habit that finds the success after a run of failures, wherever it happens.

The fixed search returns two rows, failure and success. Two rows are enough to change the finding from an attempt to a compromise.

The broken search's filter, action=failure, was a natural first step and the wrong final one. Filters chosen for the question that started a search often hide the answer to the question that matters.

The fix counts by outcome instead of filtering to failures. A search that removes the successes by design can never show the one that matters, which is why Module 2's filters are taught with what they leave out.

Module 1 also teaches the habit this module began: reading a search someone else wrote, step by step, before trusting it. Every later module returns to it, and Module 10 makes it a method.

02

Combining, Unpacking, Reusing

Modules 5 to 7

Modules 5 to 7 follow a question across sources, take apart text and lists, and turn searches into tools a team shares.

index=o365 user="svc-netops@ne.com"
| stats count min(_time) as first max(_time) as last

The search names the user in quotes because the value contains a dash and an at sign. Quoting values with punctuation is a small habit that avoids a class of surprises, which Module 2 explains.

svc-netops is a service account, named for network operations, and service accounts do not usually browse SharePoint. A familiar name doing an unfamiliar thing is a pattern the course meets more than once.

min and max of _time give the burst's start and end in one stats. Eighteen minutes from first to last, for 146 actions, is about eight a minute, sustained.

The Office 365 log records each file action with the user, the file's path and the operation. 146 of them in eighteen minutes is faster than anyone reading files would work, the pace of a script or a sync.

The first action came at 16:03, one minute after the account's second sign-in of the month at 16:02, a sequence Section 9.4 found from the other side. Two searches on two sources, joined by time, tell the story.

146 actions in eighteen minutes by an account that signs in twice in the month. Connecting it to the account's two sign-ins means combining sources, Module 5's subject; reading what it took means taking file paths apart, Module 6's.

Module 7's knowledge objects matter more than their place in the middle suggests. A team's macros, lookups and event types are how searches survive the person who wrote them, and how a detection written once runs everywhere.

03

Time and Models

Modules 8 and 9

Module 8 puts events in time: bins, baselines, running windows, sessions. Module 9 searches data models with tstats and checks what they cover.

index=endpoint host=NE-BENNETT-LT EventCode=1
    (process_name=notlocker.exe OR process_name=vssadmin.exe
     OR process_name=wbadmin.exe OR process_name=bcdedit.exe)
| stats count min(_time) as first by process_name
| sort first

NE-BENNETT-LT is named for its user, the company's convention for laptops. A machine name that leads to a person is the start of the next question: who was using it, and what else they did that morning.

The search names four programs in a base search with OR and parentheses, the Boolean logic Module 2 teaches. Without the parentheses, the OR would bind differently and the search would read far more than one laptop's events.

bcdedit ran twice, the only program of the four to run more than once. The count column records it without explanation, and the command lines, in the next search, would say what each run changed.

vssadmin, wbadmin and bcdedit are ordinary Windows tools, used here for an unordinary purpose. A search for their names alone would find administrators; the sequence, on one laptop, in three minutes, after an unknown program, is what makes it an attack.

sort first ordered the four programs by when each first ran, which turns a count into a sequence. Section 8.3 found the same sequence starting six minutes after an encoded download on the same laptop.

Four programs in under three minutes, ending a laptop's ability to recover. Module 8 finds the machine by comparing it with the rest of the fleet; here, knowing the programs, one search shows the sequence.

index=web sourcetype=squid:access url=*sync-telemetry-cdn.net*
| stats count sum(bytes_out) as bytes max(bytes_out) as largest

max(bytes_out) picked out the one large request among the 25, the statistic that turns a sum into a shape. Sum and max together say whether a total is many small things or one large one.

The 268 MB upload came about half an hour after the last of 24 small check-ins every ten minutes, as Section 8.8 found. A pause before a large transfer is common: the program gathers, then sends.

url=sync-telemetry-cdn.net matched the domain anywhere in the address, a wildcard term Module 2 teaches. Leading wildcards are slow on large data, as Module 10 notes; on a month of proxy logs they are fine.

The domain's name, sync-telemetry-cdn, is chosen to look ordinary, like a vendor's update service. Names cannot be trusted; the pattern of the traffic, regular and then one large burst, is what gives it away.

The largest single request, 268,435,456 bytes, is exactly 256 MB, a round size that software produces and people do not. The other 24 requests together carry under ten thousand bytes, the small check-ins of a program reporting in.

25 requests to one domain, one of them carrying almost all the data. Module 8 times the check-ins and the upload that followed them.

Module 9's data models are how most published detections are written, and its findings about what a model leaves out are some of the course's most practical: a detection is only as good as the data its model holds.

04

Optimization and Validation

Module 10

Module 10 asks of every search what it costs and whether its answer is right: how a search runs, filtering first, limits that cut results short, public detections, checking answers, checking AI-written SPL, and SPL2.

index=network sourcetype=pan:traffic dest_ip=45.137.21.88
| stats count sum(bytes_out) as bytes_out

The 35 connections are the firewall's view; the endpoint's Sysmon records the same connections from the laptop's side, which Module 9 finds when the data model counts them twice.

dest_ip in the base search named the address exactly, so the search read only its 35 connections. When an investigation has an address, naming it in the base search is the cheapest way to see everything that touched it.

The transfer's dates overlap the SharePoint downloads of 12 March. Two findings on the same afternoon, one taking files and one sending data out, are the kind of connection a timeline across sources is built to make.

45.137.21.88 appears in the firewall's log only between 14:07 on 12 March and 15:45 on 14 March, after the attacks had begun. A destination seen only during an attack is one of the clearest signs a search can find, and Module 8's first-and-last searches are built to find such things.

The firewall's events carry the byte counts that make this visible. A source that records how much was sent, not only whether a connection happened, is what any question about data leaving needs.

35 connections and 432 MB. Module 9 finds that the traffic data model cannot see the bytes; Module 10 finds that a search for the largest uploads ranks this third, behind backups. Both are findings about searches as much as about the attack.

The exercise reproduces a figure the course reaches again in Modules 9 and 10, each time by a different route, and each time it agrees.

sum(bytes_out) added the bytes of every connection to the address. count beside it shows how many connections carried them, the two numbers any question about data leaving needs together.

432,035,449 bytes is about 412 MB, and nearly all of it went in one transfer, which Section 9.7 found. A sum hides the shape of what it adds up; a breakdown by connection would show the one large transfer among small ones.

index=network sourcetype=pan:traffic dest_ip=45.137.21.88
| sort - bytes_out
| head 3
| table _time bytes_out

The large transfer came at 15:30, the small ones around it are the connection's own traffic. A breakdown of a total, here a sort, is what turns a number into a description of what happened.

The completion exercise is sum(bytes_out). The number it produces is the month's most consequential, and the course reaches it several ways, which is how it shows that the number is right.

Module 10 is where the course turns on itself. It takes searches from earlier modules, from Splunk's published detections and from AI assistants, and checks each against the month, finding where each would have misled.

05

The Guess That Worked

Reading a finding to its end

A finding is not complete until it says what happened next. The guessing's 83 failures matter because of what came after them.

The search ends with table, not stats, because the question is about one event and its details. Choosing between a table of events and a table of counts is one of the first decisions in any search.

The sign-in was to Exchange Online, the mailbox, the same application the 83 failures had targeted. An attacker who guesses a password usually wants something specific, and the application says what.

action=success in the base search kept only successes from that address, one event. Finding the one success among 84 events from an attacker's address is the moment an investigation of failures becomes an investigation of a compromise.

r.scott at 22:55 on 2 March, signed in to Exchange Online from the address that had failed 83 times. Section 0.5 found the guessing; this search finds its result. A finding that stops at the failures misses the one event that made them matter.

Following a finding to its end is a habit, not a command. The searches that do it are simple; what matters is asking the next question, what happened after, every time.

06

What You Will Be Able to Do

Four abilities

The modules build four abilities, each on the one before.

What you will be able to do
1Answer a question from events
any sourcetype, any fieldModules 1 to 4
2Follow it across sources and text
combine, unpack, reuseModules 5 to 7
3Put it in time and across models
timelines, baselines, tstatsModules 8 and 9
4Defend the answer
cost, limits, checksModule 10
Four abilities, each built on the one before, and each tested on the same month.

Each ability is tested on the attacks in the timeline. By the end of Module 10, most will have been found with more than one search, and the findings checked against each other.

The first ability is enough for most daily questions. The second and third are what investigations need, when an answer lies across sources or in a sequence. The fourth is what makes any of them safe to act on.

The fourth ability is the one that makes the others worth having. An answer that cannot be defended, by a second search, a check of its parts, or a known case it finds, is not yet an answer.

The abilities are tested throughout: by predictions before searches, by exercises that start from a question or a broken search, and by each module's questions that no section answered.

07

What the Course Is Not

Its boundaries

The course teaches the language, not the product's menus, which change between releases. It teaches how to build and check detections, not a catalog of them. And it stays at the search bar: indexes, inputs, clusters and apps are administration, another role's work.

What the course is not
A product tour
Menus change; searches last.It teaches the language
A detection catalog
Searches without reasons do not transfer.It teaches how to build and check them
An administration course
Indexes, inputs and clusters are another job.It stays at the search bar
Each is useful, and each belongs somewhere else.

The third boundary is a matter of roles. Analysts search the data administrators collect; the course gives analysts what they need to ask administrators for the right things, without teaching the administration itself.

The first boundary matters because Splunk's interface changes between releases and between Splunk Enterprise and Splunk Cloud Platform, as Section 0.4 found. The language underneath changes far more slowly, so the course teaches what will still be true next year.

Each boundary keeps the course focused on what transfers. A search written well in SPL works in any Splunk deployment that holds the data; the habits that check it work everywhere.

Within those boundaries, the course goes deep. It covers the SPL an analyst needs from the first search to the checks a senior analyst runs before trusting a detection, and it shows each on data where the right answer is known.

08

Starting

Module 1

Module 1 begins at the bottom rung: turning a question into SPL, the search pipeline, how commands are ordered, how to learn a sourcetype you have never seen, field values and types, missing fields, checking a result against its events, and reading a search you did not write. Every one of those is used in every module after it.

index=azuread category=SignInLogs
| stats count(eval(action="failure")) as failures count by user
| where failures > 0
| stats count as accounts_with_failures sum(failures) as failures

The where kept accounts with at least one failure before the second stats counted them. A condition between two statistics is common, and its place, after the first stats, is the only place it can go.

The search counted failures inside stats, then counted the accounts with any. Two stats in a row, the second summarizing the first, is a shape the course uses from Module 4 onwards.

Most accounts failed at least once in a month, which is ordinary: people mistype. The course's work is finding the few failures that are not ordinary among the many that are.

The habits are four, and they are the course in brief. Ask questions you can check. Build searches one step at a time, counting as you go. Follow findings to their end, success included. And defend every answer.

Practice

Three questions about the month's attacks, each answered by one search.

Each answer is the size of an attack: prompts, connections, requests. Rerun each search, and keep the numbers; they come back in the modules that find these attacks properly.

Module 0 ends here. Module 1 begins the course.