Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
0.3 The Sourcetypes a Security Analyst Searches
Introduction
An attack leaves traces in several places: a sign-in in the identity provider, a process on a laptop, a connection through the firewall, a file downloaded from the cloud. Each place writes its own kind of record, and Splunk stores each kind under its own sourcetype.
Knowing which sourcetype records what is half of knowing where to search; the other half is knowing what each calls things. This section introduces the sourcetypes this course searches, one family at a time, with a count of each.
You will split Entra ID's single sourcetype into four kinds of record, group the Windows Security log and Sysmon by event code, see what the firewall and the VPN gateway do with connections, read a proxy request as fields, find the program on the Linux servers that records login attempts, and count five more sources from the cloud, mail and DNS.
By the end you should be able to say which sourcetype answers a given question, and find the field in it that says what kind of event each record is.
The diagram groups the month's sources into eight families, in roughly the order an attack crosses them. The top row is the core of the course; the bottom row comes in from Module 5 onwards, as investigations reach beyond one source.
Sourcetypes in this course
The sample month, counted 3 October 2026The card leaves out some of the month's smaller sources: Azure's resource and activity logs, Defender's incidents and alerts, Windows DNS server logs. They appear where a question needs them, and the ladder later in this section is enough to learn each.
The sourcetype names come from the Splunk add-ons that collect each kind of data, and the same names appear in most Splunk deployments that collect the same products. A search written for azure:monitor:aad here is close to one that would work elsewhere, which is part of why the course uses real sourcetype names.
The card is a reference to keep beside the lab while the course runs. Each line names a sourcetype and what it records, and most name the field that divides it into kinds of event, which is always the first field worth grouping by.
Identity
azure:monitor:aadEntra ID, Microsoft's identity service, records every sign-in to the company's cloud services and every change to its directory.
index=azuread
| stats count by category
All four together are 23,843, the index total Section 0.1 counted. The kind field divides the whole source, so its counts always add up to the total.
stats count by category is the second rung of this section's ladder applied to Entra ID: the source's own field that says what kind each event is, grouped and counted.
The non-interactive sign-ins outnumber the interactive ones by more than one and a half to one. They are token refreshes by applications acting for a user, and a count of sign-ins that includes them describes applications as much as people.
Four kinds of record under one sourcetype. Most of this course's identity searches read the interactive sign-ins, category=SignInLogs, the ones a person typed a password for.
The 532 audit events are the smallest category and among the most important: every change to an account, a group or an application's permissions is one of them. Module 6 reads one, an application granted access to mail.
The four categories are Microsoft's own names for its log streams, and the same names appear in Microsoft's other tools. Learning them once serves wherever Entra ID data is searched.
Audit events have different fields from sign-ins: an operation, a target, an initiator. A search that mixes the two categories mixes two shapes of record, and fields present in one are empty in the other.
The completion exercise is category. Any search on this sourcetype that does not name a category reads all four kinds, which is rarely what a question about sign-ins means.
Entra ID is a structured source: its events arrive as records with named fields rather than lines of text. Searching it means knowing its field names, user, src_ip, action, app, which the course introduces as each is needed.
Windows
WinEventLog:SecurityThe Windows Security log records logons, Kerberos tickets and changes to Active Directory, on the domain controllers and other machines that send it.
index=wineventlog sourcetype=WinEventLog:Security
| stats count by EventCode
| sort - count
| head 6
4768 and 4769 together, 737 events, are Kerberos at work: a ticket to prove who a user is, then tickets for each service they reach. Attacks on Kerberos show as unusual patterns in these two codes.
This log comes mostly from the company's two domain controllers, NE-DC01 and NE-DC02, which see logons to domain accounts, with smaller shares from servers and laptops. Section 9.1 found some of its non-logon events mapped into the Authentication data model, a reminder that a source's own event codes are the reliable way to select its events.
The EventCode values are strings of digits, and a search can name them as numbers or text in the base search: EventCode=4624 matches either way.
4662 and 5136 are directory changes: an object accessed and an object modified. Small in number, they are the events an investigation of changes to accounts or groups reads first.
Logons and Kerberos dominate. Each event code is a different kind of record with its own fields, so a search on this log almost always starts by naming the codes it wants.
Windows event codes are documented by Microsoft, and an analyst learns the common ones by heart over time: 4624 for a logon, 4625 for a failure, 4768 and 4769 for Kerberos. The course introduces each where it is first used.
Endpoint
SysmonSysmon, a Microsoft tool installed on Windows computers, records what runs on each one: processes, network connections, files, registry changes.
index=endpoint
| stats count by EventCode
| sort - count
NE-BENNETT-LT, a laptop that appears in the Windows log too, is where Section 8.3 found the ransomware's preparation in Sysmon's process events. A machine's Sysmon events are the most complete record of what ran on it.
Sysmon is not installed everywhere. Section 9.2 found three machines that send process starts and nothing else, and a source's coverage across machines is as much a part of knowing it as its fields.
Event 8, a thread started in another process, has 44 events, the fewest. Rare event codes are often the interesting ones, because what an attacker does is rarely what a computer does every day.
Process starts lead. Sysmon's process events record the full command line of everything that ran, which is where the course finds encoded PowerShell and the ransomware's preparations.
Sysmon and the Windows Security log both number their events, and the numbers mean different things in each: Sysmon's 1 is a process start, Windows's 4688 would be. An event code means nothing without its sourcetype.
Network
pan:traffic and cisco:asaTwo network devices sit at the company's edge: the firewall that every connection crosses, and the VPN gateway that remote staff connect through.
index=network sourcetype=pan:traffic
| stats count by action
The firewall's events have no host in the lab, as Section 0.2 found, and no _raw either; their fields are all a search has. That makes the field list, which Module 1 teaches how to read, the only map of this source.
The 440 denied connections are the firewall doing its job, blocking traffic its rules do not permit. Allowed connections, the large majority, are where an attacker's traffic hides, because it was allowed.
The firewall's action field has five values, more than allow and deny, because its log also records threat alerts and some session outcomes. Listing a field's values before filtering on it is how the full set is found.
The gateway records something different: people connecting in.
index=network sourcetype=cisco:asa
| stats count by action
The gateway's 532 events are few beside the firewall's 2,095, because it records logins and sessions, not every packet. Small sources can carry large findings; the VPN spray is in these 532.
The gateway's events are lines of text, like syslog and the proxy, and the action field is extracted from each line's message. A typical line names the VPN group, the user, the address assigned and what happened, such as a session established for r.scott.
The gateway's three values, success, allowed and failure, describe a login and then the session it opens. The 43 failures are rejected logins, which the course reads more than once.
Two sourcetypes, two vocabularies. The firewall writes allow for a permitted connection; the gateway writes allowed. Any search across both has to use each one's own word, or Module 9's data models, which give them one.
244 allowed sessions and 245 successful logins are nearly equal, as they should be: each successful login opens a session. A pair of counts that should match, and do, is a small sign that a source is being read correctly.
The broken search returned 0, a count, not an error. Every value a search names is worth checking against a count by that field in that source, the fourth rung of the ladder.
The fix uses the gateway's own value. A value copied from another source matches nothing, the same silent miss as a misspelled field name.
The first mistake shows up as numbers that are too large: sign-ins counted with token refreshes and audit events. A count far above what the company's size suggests is the clue to look for a kind field.
The third mistake is the one that wastes most time: a search that runs well on the wrong source looks like proof that something did not happen. Asking which source would record the event, before searching, is the habit that prevents it.
All three are about matching a question to a source and a source to its own words. Each is avoided by the ladder below, run once on any source before trusting a search on it.
The second mistake is this section's own finding, and the first is Section 1's. The third is avoided by the card: logons are in the Windows Security log and Entra ID, not in Sysmon.
Network devices see traffic, not people. The firewall's events name addresses and ports; the gateway's name the account that logged in. Connecting an address to a person usually needs a second source, the gateway's logins or the identity service's sign-ins.
Web
squid:accessThe proxy handles every web request from inside the company. Its events are lines of text, from which Splunk extracts the fields at search time.
index=web sourcetype=squid:access
| head 1
| table src_ip url status bytes_out
Status 200 means success, and every one of the proxy's 7,190 requests in the month has it. A field whose values never vary cannot separate anything, which is worth knowing before writing a search that filters on it.
The fields come from the line by the rules of the proxy's sourcetype. table chose four of them; the proxy's events have more, including the method and the content type, which the raw line showed.
src_ip is an internal address, 10.0.1.112, a machine inside the company. The proxy sees machines, not people, so a question about which person visited a site needs the machine's user from another source.
bytes_out is a byte count the proxy recorded for the request. Field names like it are the add-on's choice, and Module 1 shows how to find what each field means in a source you have not seen.
The same request as Section 0.2's raw line, now as fields. The company's websites have their own log, access_combined, which records visitors coming in rather than staff going out.
The proxy and the firewall both see outbound traffic, at different layers: the proxy sees web requests by name, the firewall every connection by address. The course finds data leaving through both: a large upload through the proxy in Module 8, and a larger transfer through the firewall in Module 9.
Linux
syslogThe eight Linux servers send their system logs, one line for each thing a program reports. The process field says which program wrote each line.
index=linux
| stats count by process
| sort - count
| head 5
The 338 sudo events are commands run as administrator. On a well-run server they are routine, made by a few known accounts; an unfamiliar account among them is worth a look.
The process field is extracted from each syslog line, the program name before the colon. Section 0.2's raw line showed one: CROND[28272], the program and its process number.
crond and systemd are routine: scheduled jobs and services starting and stopping. sudo records commands run with administrator rights, a small source with a large share of what matters on a server.
sshd first, the program that accepts SSH logins. Most of the server attacks in this month arrive as SSH login attempts from addresses outside the company.
The eight servers' sshd counts add up to 2,411, the total the previous search found: host divides sshd's events exactly, as category divided Entra ID's.
process=sshd went in the base search as a field=value term, so the search read only sshd's events. On a server's full log, that one term removes crond, systemd and every other program before counting.
The counts are almost equal, about 320 for each of the busiest servers. Unlike the count of all events in Section 0.2, where the web servers led, SSH activity reaches every server alike, which the course explains when it finds where the attempts come from.
The write exercise counts sshd per server. Grouping a source's events by its kind field, then by host, is the course's standard first look at any source on machines.
syslog is the least structured source in the course. Its events are lines of free text, and fields beyond the program, host and time often have to be taken apart with rex, Module 6's subject.
Cloud, Mail and DNS
The rest of the monthFive more sources fill in the rest of the company: its Office 365 files and mailboxes, its AWS account, its mail as Defender sees it, its name lookups, and its own websites.
| tstats count where index=* by index sourcetype
DNS events record the names machines looked up before connecting anywhere. A lookup of an unfamiliar domain is often the earliest trace of a connection that other sources record only later, which is why DNS is worth searching alongside the firewall and the proxy.
AWS CloudTrail is the largest of the five, though the company runs most of its work on Microsoft's cloud. Volume says little about importance; CloudTrail records every API call, most of them automated.
tstats counted all five in one search, as in Section 0.1. A count by index and sourcetype across everything is the quickest first look at any deployment, the first rung of Section 0.2's ladder.
The Defender mail events are one of several kinds Defender sends; the lab holds more than a dozen of them under the one sourcetype ms:defender:eventhub, 66,049 events in all, and a search names the kind it wants.
Each is large, between 8,852 and 16,880 events. They are where an investigation goes once it leaves the endpoint or the identity service: files taken from SharePoint, which this month's svc-netops account did, calls to a cloud provider's API, a message delivered, a domain looked up.
The first rung, a count, also says how much work searches on the source will be. A source of hundreds of events can be searched any way; one of millions rewards the care Module 10 teaches.
The third rung, one event of each kind, is where field names come from. A field seen in a real event is one a search can rely on; a field remembered from another deployment may not exist here.
The fourth rung records the source's vocabulary. Writing down that the gateway says allowed and the firewall allow takes a minute and saves every later search on either from a silent miss.
The ladder is how to learn any of these, or any sourcetype at work. The second rung matters most: every source has a field that says what kind of event each record is, and grouping by it is the map of the source.
No single source tells the whole story of an attack. The month's password guessing shows in Entra ID; its spray shows in Entra ID and the VPN gateway; its ransomware preparation shows in Sysmon; its exfiltration in the firewall and the proxy. The course reads each where it shows most clearly.
Sources in the Rest of the Course
Where they come backModule 1 teaches how to learn a sourcetype you have never seen, the ladder above in full. Module 5 combines sources, and Module 9 shows how data models give many sources one shared set of names and values. Section 0.4 turns from the data to the platforms that hold it.
The habits are four. Count a new source before searching it. Find its kind field and group by it. Look at one event of each kind, and note its fields. And use the values the source writes, not the ones another source, or memory, suggests.
Practice
Three questions about the sources, each answered by a count by its kind field.
All three answers come from grouping a source by its kind field. Each of the three is the first number worth knowing about its source.
Next, Section 0.4 compares Splunk Enterprise, Splunk Cloud Platform and SPL2.