Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
0.2 How Splunk Stores Events
Introduction
Splunk stores data as events, and every event carries the same small set of fields whatever produced it: its original text, its time, the index it is stored in, its sourcetype, the host that sent it, and its source.
Everything else a search uses, an account, an address, an outcome, is extracted from the event when the search runs. Knowing the default fields is knowing where any data is and how to ask for it, in any deployment, whatever its sources. This section looks at the events themselves, one default field at a time.
You will read one Linux event with its default fields, see a proxy log line before any field is extracted, check that _time came from the timestamp in the text, find two sourcetypes sharing one index, count events by the machine that sent them, find a field the lab does not record, and meet the newest event in the month and the lab's idea of now.
By the end you should be able to say, for any event, where it lives, what format it is in, when it happened and where it came from.
The diagram shows one event, the Linux line from Section 1, and the six fields every event has. The bottom row is the larger set extracted at search time, which depends on the event's sourcetype and is the subject of Module 1.
The default fields
Splunk Search Manual and Knowledge Manager Manual, read 3 October 2026The underscore at the start of _raw and _time marks them as internal fields, ones Splunk manages itself. Fields with a leading underscore are hidden from some commands' default output, which is why a table usually names them explicitly.
Four of the six, index, sourcetype, host and source, are also indexed fields: Splunk records them in its index files, so searches and the tstats command can use them without reading the events themselves. Module 9 builds on that.
The card lists the six default fields. Splunk's documentation calls them default fields because Splunk adds them to every event when it is stored, whatever the data looks like.
One Event
The default fieldsThe quickest way to understand stored data is to look at one event and its default fields.
index=linux sourcetype=syslog
| head 1
| table _time index sourcetype host _raw
WEB-NGE-MCR-02 is one of the company's two web servers. m.webb, whose session closed, is the security architect, so this event is someone with administrative access finishing work on a public-facing machine.
The five fields shown are the ones every source shares. Asking for them by name with table works on any event in any index, which makes this search a template for a first look at any source.
The event's text is short, as Linux log lines often are: what happened, and to which account. The account name, m.webb, is in the text, not in a field of its own, until a search extracts it.
A session closing on a web server. Five of the six default fields are in the table; the sixth, source, is the subject of Section 7.
The sort after stats puts the busiest server first, which is the order most questions about machines want.
A by clause on host gives one row per machine; adding sourcetype, by host sourcetype, would give one row per machine and format, a finer map that matters when a machine sends several kinds of data.
The eight host names follow the company's naming convention: SRV for servers, WEB for web servers, NGE for the company, MCR and BRS for two of its sites. Names like these are worth learning early, because they are how an analyst recognizes a machine in a result.
The completion exercise is host. It is the default field analysts use most after index and sourcetype, because so many questions are about one machine.
The newest Linux event is from 12:46 on 15 March, though the other sources' data ends just before noon. Section 8 comes back to it; for now, it is a reminder that head shows whatever is newest, and newest is not always what was expected.
_raw
The original text_raw is the event as it arrived, before any field was taken from it. For a log line, it is the line.
index=web sourcetype=squid:access
| head 1
| table _raw
TCP_MISS/200 is the proxy's way of saying it fetched the page fresh and the site answered with status 200, success. Codes like these are what a search for blocked or failed requests would filter on, once extracted.
The address on the line, cdn.jsdelivr.net, is a public content network that ordinary web pages load scripts from. Most proxy lines are like this one: routine, and the reason a search has to say exactly what it is looking for.
The line follows the proxy's own format, which its documentation defines and its sourcetype in Splunk knows. Reading a few raw lines of any new source is the quickest way to learn what is in it, before deciding which fields to search on.
A single line of text. The proxy wrote it; Splunk stored it unchanged. Fields such as the client address or the status are not stored separately for this kind of data; a search extracts them from _raw each time it runs, by rules attached to the sourcetype.
Not every source arrives as a line of text. Many security sources, Entra ID sign-ins and Defender events among them, arrive as structured records whose fields are named in the data itself. In the lab, those sources are held as fields with no _raw, another difference Section 0.6 lists.
_time
When it happenedEvery event has a time, however its source writes it. Splunk reads it from the event when the event is stored and keeps it in _time.
index=web sourcetype=squid:access
| head 1
| rex "^(?<raw_epoch>\d+)\."
| eval time_epoch=strftime(_time, "%s")
| table _time raw_epoch time_epoch
The number 1773575758 is the time 11:55:58 on 15 March 2026, minutes before the month's data ends. Unix time is how many systems store times internally, and SPL can read and write it directly.
The fraction after the period, .510, is milliseconds, which _time keeps too: Splunk stores times to the millisecond and beyond, and the lab shows them in its time column.
rex took the digits before the first period from _raw, and eval wrote _time back as seconds with strftime. The two numbers being equal is a small proof of where _time came from; Module 6 teaches rex and Module 3 teaches eval and strftime.
1773575758 in both. The proxy writes its time as seconds since 1970 at the start of each line, and Splunk took _time from there. Other sources write dates in other forms, and the sourcetype's settings tell Splunk how to read each one.
_time is what every time range in SPL tests, from earliest and latest to the time picker. An event whose timestamp was read wrongly when it was stored is found at the wrong time by every search, which is why sourcetype settings for timestamps matter so much to the people who run Splunk.
index
Where events are storedAn index is a store of events, kept on disk and searched as a unit. Organizations divide data among indexes by kind, by retention or by who may see it.
index=network
| stats count by sourcetype
The firewall's events outnumber the gateway's four to one. Searching the index without naming a sourcetype would put the gateway's logins in a minority of every result, easy to overlook.
stats count by sourcetype is the second rung of this section's ladder. Run on any index, it shows how many formats a search on that index will meet, and how much of each.
The firewall logs connections, allowed or denied; the gateway logs VPN logins and sessions. A count of everything in the network index mixes the two, which is rarely what a question means.
Two sourcetypes in one index. A search that names only index=network reads both, firewall connections and VPN logins together; naming the sourcetype as well reads one.
The broken search returned a count of 0 rather than no rows at all, because stats count always returns a row. A zero is still a result, and the habit of asking why it is zero is what the fix exercise practices.
The gateway's 532 events are the VPN logins this course returns to several times, including the spray on the morning of 5 March. Finding them starts with knowing they are in network, not in an identity index.
The broken search's sourcetype was right and its index wrong, and the combination matched nothing. Index and sourcetype together narrow a search best when both are right; one wrong makes the other irrelevant.
The fix names the right index. A search in the wrong index does not fail; it reads events that are not the ones wanted and finds none of the ones that are.
The lab's month has eleven indexes, named for their sources: azuread, network, web, linux, endpoint, wineventlog, dns, defender, o365, aws and azure. Real deployments choose their own names, and the first search on a new deployment is usually the one that lists them.
sourcetype
What format the data is inA sourcetype names a format. It decides how Splunk reads the timestamp and which fields a search can extract, so two sources with one sourcetype are read the same way.
The proxy's 7,190 events outnumber the network index's two sourcetypes together, 2,627. Size and number of formats are different questions, and a map answers both.
The three rows are a small version of Section 0.1's map of seven sources. Every course search that reads more than one source starts from a map like it.
The parentheses around the OR keep the two index terms together. Without them, an OR between terms in a base search binds the way Module 2 explains, and a search can read more or less than meant.
The write exercise maps two indexes: the firewall and the gateway in network, the proxy in web. A search over several indexes with OR is how the course reads more than one source at once, from Module 2 onwards.
None of the three raises an error. The defense is the same each time: count by index, sourcetype and time before trusting a search's silence.
The first mistake is the commonest when moving between deployments, where the same data can sit in differently named indexes. The second makes counts mean less than they seem, two sources added together as one.
The third mistake matters for alerts. A scheduled search that ends at now and runs every hour reads each hour once; an event stamped in the future is read by none of them until its time arrives, and one stamped far in the past is never read at all, because no later window reaches back to it.
All three mistakes find less than is there. A count by index and sourcetype before any question is the habit that prevents the first two.
Sourcetype names follow conventions set by the add-ons that collect the data: cisco:asa for the gateway, pan:traffic for the firewall's traffic log, azure:monitor:aad for Entra ID. Section 0.3 introduces the sourcetypes this course searches and what each records.
host
Which machinehost names the machine that produced the event: a server, a firewall, a laptop, a gateway.
index=linux sourcetype=syslog
| stats count by host
| sort - count
SRV-NGE-BRS-DB01 is the only server at the second site, with 642 events, about the same as its peers. Hosts from different sites behaving alike is what a healthy estate looks like in a count by host.
The two web servers face the internet, and the scanners and attacks that come with that are part of why they are busiest. The course later finds a thousand failed SSH logins spread across all eight servers.
The sort put the busiest first, as in Section 0.1. A count by host sorted the other way, quietest first, is how an analyst finds a machine that has stopped sending events, which can mean a broken collector or an attacker who turned logging off.
The counts are close together, between 606 and 861, because the servers do similar work every day. A server whose count suddenly doubled, or dropped to nothing, would be the first thing such a search shows.
Eight servers, the whole Linux estate. The two web servers are the busiest, which suits machines that face the internet. A count by host is the first look at any group of machines, and it shows at once whether one is unusually quiet or loud.
The second and third rungs are where the default fields earn their place: index, sourcetype and host are present on every event in a real deployment, so the same three searches map any source without knowing its own fields.
The first rung uses tstats, as Section 0.1's map did, because it counts every index quickly without reading events. The other rungs use stats on one index at a time.
The fourth rung is the one this section has been doing: one event, its default fields, its raw text. Looking at a single event before writing a search on a source is the most reliable way to get its field names right.
The ladder is four searches for any deployment, in this lab or at work. The first two make a map; the last two show what is on it.
host is not always the machine an event is about. The VPN gateway's 532 events all have the gateway, ASA-NGE-MCR-01, as host, though they describe logins by people working elsewhere; the addresses and accounts inside the events say who.
In the lab, the firewall's events carry no host at all, one more difference for Section 0.6. Knowing which machine a host field names, the sender or the subject, is part of learning a sourcetype.
source
Where the event came fromsource names the file, stream or input an event came from: a log file's path, for example. In a real deployment it is present on every event.
index=linux
| stats count(source) as with_source count
A field the lab lacks is not a field a search can test; a condition on source here would match nothing, as a misspelled field would. Section 0.6 lists every such gap so none of them is learned as a habit.
The one Windows log that does carry source in the lab is the System log, whose 443 events all have Service Control Manager there: in this data it names the Windows component that wrote the event, not a file. It is a small source the course reads rarely.
In your own Splunk, the same search will show a source on every event, and its values are worth a look: they say which files and inputs feed each sourcetype.
count(source) counts the events that have a value for the field, 0 here, beside count, all 5,332. The same pair is how any field's coverage is checked, a habit Section 10.8 uses on searches written by others.
None of the Linux events have it in the lab. The lab's month was built from data that does not keep source for most sourcetypes; only one Windows log has it. Searches in this course therefore never filter on source, and Section 0.6 lists this with the lab's other differences from Splunk.
In a real deployment, source is useful for telling apart events of the same sourcetype from different files: an application's access log from its error log, for instance. Searches written for real deployments often filter on it, and those filters have no effect in the lab.
Time and Now
Where the month endsA search reads events within a time range, and a range has two ends. In Splunk, the start is included and the end is not, so hourly searches from one hour to the next never read an event twice.
In the lab, now is fixed to the time of the newest event a search reads, so that a relative range such as the last day means the same thing every time the lab runs.
index=linux
| stats count as all_time
| appendcols [search index=linux latest=now | stats count as up_to_now]
The course's searches mostly have no time range at all, because the lab's month is small and fixed. In a real deployment, every search should have one, and the time range is the first thing to check when two people get different counts.
latest=now sets the end of the search's time range in the base search; Module 2 covers earliest, latest and the time picker that sets them for you in the Search app.
The newest Linux event itself is ordinary, a session closing. Its time is what makes it stand out: about three-quarters of an hour after every other source's data ends. In a real deployment, an event stamped later than everything else would point at a server clock running fast, worth fixing because every search's time range depends on it.
appendcols put the two counts side by side, one from a search with no end time and one ending at now. Comparing a count with and without a condition, one change at a time, is how the course finds what each condition does.
5,332 and 5,331. For a search over the Linux events, the lab's now is 12:46:26 on 15 March, the newest event's own time, and the range ending at now leaves that one event out.
In Splunk, now is the clock, and Section 10.9 noted that All Time in SPL has no end at all, so it includes even events stamped in the future. Clocks on real machines drift, and such events are not as rare as they sound.
Practice
Three questions about where the month's events are stored, and when.
The third question is the one event a range ending at now leaves out. Run both counts yourself, and notice that a search ending at now is the default for most scheduled searches, which is why an event with a wrong clock can be missed by every alert.
Next, Section 0.3 introduces the sourcetypes this course searches, and what each records.