Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
How macOS Endpoint Investigation Is Structured: Questions, Phases, Two OS Versions
Introduction
The evidence on a Mac does not decay at one rate, and the fastest of it is gone in hours.
That single fact sets the order of everything in this course. A syllabus that taught the durable artifacts first would be teaching you to arrive after the volatile ones had already rolled off, which is a lesson nobody needs twice.
So this section is about lifespan: which stores disappear, how fast, and why the number everybody quotes for the log store is not a number at all.
Scenario
An incident on NE-VANCE-MBP is suspected to have started eleven days ago. The collection plan says the unified log will cover it, on the basis that macOS keeps thirty days. The plan is signed off, the collection is scheduled for the following week, and nobody checks the machine.
A size cap, not a time limit
Which is the whole misunderstandingStart with how the log store decides what to throw away, because it is not what the retention question assumes. Published research into the store describes the mechanism: the entries that survive to permanent storage are kept in files under a persistent diagnostics directory, and the logging daemon maintains them according to the total size of that folder rather than the age of their contents.
HOW THE LOG STORE DECIDES WHAT TO DROP
the rule keep the folder near a target total size
the target roughly 520 to 530 MB
each file up to about 10.5 MB
so the folder holds on the order of 50 files
what it is not a number of days
Four mechanical facts and not one of them mentions time. Published research adds that there does not appear to be any way to set the size allowance for these files, so this is not a policy an organization can raise by configuration or buy its way out of with more disk.
The consequence is the thing worth carrying: how far back the log reaches is decided by how fast the machine fills that space. Retention is an outcome of behavior rather than a setting, which is why it varies between two machines that were imaged from the same build on the same day.
It also explains a common and wasted argument. When somebody says the log should have covered the period, they are usually reasoning from a policy they believe exists, and there is no policy to point at. The store behaves the way a fixed-size buffer behaves, which is to say it discards the oldest thing whenever it needs room, and nothing in the operating system offers to hold more.
The one lever an organization does have is elsewhere, and it is worth knowing so the conversation goes somewhere useful. Shipping entries off the endpoint to a collector removes the size cap from the equation entirely, because retention then belongs to whoever runs the collector rather than to a fixed folder on a laptop. That is an estate decision with a cost attached, and it is the honest answer to anybody asking how to get more than a few days.
So the window is a property of the machine
And it has been shrinkingWhat the cap works out to in days is a question published observation answers with a trend rather than a figure. When unified logging was introduced it was not unusual for a healthy machine to keep log files for up to twenty days, and as the quantity of entries has steadily increased the period covered has reduced.
That direction of travel matters more than any single number, because it means a figure learned a few releases ago is not merely imprecise but wrong in a predictable direction. Anybody quoting a retention period from memory is quoting a machine they last looked at, on a version that has since been superseded. The failure mode is not ignorance but confidence: the figure was accurate once, which is exactly why nobody thinks to check it again.
The auditor above holds the scenario's collection plan. Several lines in it are sound and one assumption decides whether the collection is worth scheduling at all. Work it before the practice card.
The trend also has a practical implication for anything written down. A retention figure in a runbook, a collection template or a service description ages badly and silently, because nothing prompts a review when the underlying behavior shifts. If a number of that kind exists in your own documentation, it is worth replacing with an instruction to measure rather than a figure to trust. An instruction stays correct across releases; a number does not, and the number is the thing somebody will quote back at you.
The same applies to anything contractual. Where a managed service commits to investigating incidents within a stated period, that period is only meaningful against a log window nobody has measured, and the two figures are usually set by different people who never compared them. Discovering the mismatch during an engagement is considerably worse than discovering it while writing the commitment. Checking it is a single measurement on a representative machine, and it either confirms the commitment or gives you a concrete reason to change it. Either outcome is worth more than the assumption it replaces, and the measurement takes a minute on a machine somebody already owns.
A busy machine keeps less
Which inverts the usual intuitionA size-capped store has an implication that runs against instinct, and it is worth stating directly: the more a machine does, the less history it keeps. Published description states that the default rotation is approximately seven days but a busy system can rotate in under twenty-four hours, and that high-volume subsystems may roll over within hours.
A fixed cap means the rate of entries sets the number of days.
Read the two paths as the same cap filled at different speeds. The machine under investigation is frequently the one doing the most, which makes the shortest window and the most interesting history the same case.
THE SAME CAP, TWO MACHINES
design workstation builds, containers, chatty subsystems
folder fills fast window short
light-use laptop mail, browser, little else
folder fills slowly window long
settings identical on both
Nothing was configured differently between those two machines. The difference is entirely in how fast each one produces entries, which means an estate cannot be characterized by a single retention figure however carefully that figure was measured on one device.
That cuts against the way fleets are usually described. An organization will happily state that its Macs retain a week of logs, meaning somebody once checked a laptop, and the design workstations doing continuous builds are the machines least likely to match that claim and most likely to be the subject of an investigation. Where a fleet figure is needed at all, it should be a range measured across representative machines rather than a single number. Picking those machines by role rather than at random is what makes the range mean anything, because role is what drives the entry rate.
It is also worth measuring the extremes rather than the average. The machine with the shortest window is the one that will disappoint an investigation, so knowing which role that is tells you where to shorten the collection deadline, and knowing the longest tells you where an older incident might still be recoverable. An average across the fleet describes no machine anybody will actually examine.
The thirty-day figure is wrong
And published work says so plainlyThe figure in the scenario's plan is worth confronting directly rather than quietly working around. Published analysis is unambiguous: the idea that logs are always retained for thirty days is largely incorrect. Entries carry a time-to-live that varies by class, and lifetimes range from minutes to effectively indefinite.
The repeated claims, and what the documentation actually says.
Read the struck-through column as things worth unlearning rather than as a list of errors, because each one is repeated in good faith by people who have not read the documentation.
Unified log retention, as the documentation describes it
Several numbers, not oneWhat follows for a report
- Thirty days is not in the documentation, and an opposing examiner needs one command to show the window on this machine was shorter.
- Entries carrying the longer lifetimes are relatively rare and in many cases less useful for forensic purposes, so they rarely fill the gap.
- A measured window, written with the date it was measured, is a claim nobody can take apart.
The figure that matters is a property of the machine in front of you rather than of macOS, which is why it gets measured rather than quoted.
Read the rows across rather than down, because both columns describe the same store. The right-hand column is the one to put in a plan, and treating retention as a single platform number is precisely what the left-hand column does.
There is a useful exception in the same research. Persistent signpost data lasts longer, on the order of weeks to months, and is the source most likely to preserve evidence for an older incident, which makes it the thing to reach for when the log window has closed on the period you care about.
The exception is narrower than it first appears, though, and worth reading carefully before relying on it. The same research notes that entries carrying the longer time-to-live values are relatively rare and in many cases less useful for forensic purposes, so what survives longest is not a representative sample of what was there. Treating it as one produces the particular kind of wrong answer that looks thorough, because the entries are real and the reasoning over them is sound; only the assumption that they represent the period is false. A finding built only on the long-lived remainder describes the part that happened to persist rather than the period as a whole. Saying so explicitly is the difference between a narrow finding and an overstated one, and it costs a sentence.
The surviving entries also skew toward particular kinds of event rather than being spread evenly, which compounds the problem. A timeline built from them can look continuous while being made almost entirely of one subsystem's output, and that shape is easy to miss when the entries are read in sequence rather than grouped by source. Grouping by subsystem before reading is a cheap guard against it, and it takes one sort.
Not every store decays alike
Which is what sets the reading orderNot every store decays alike, and the log is only the fastest clock running. Published observation puts activity and interaction records at roughly twenty-eight to thirty days, describing the figure as undocumented and drawn from observed behavior rather than a published guarantee, which is a distinction worth preserving when the number reaches a report.
Five stores by survival, and the collection order that follows from it.
Read the bars as survival time rather than as importance, because the shortest bar is the one that decides what you do first.
The log store on the machine you are holding
Measured, not assumedWhy the record carries a date
- A window quoted without the date it was measured is an assumption, because the edge moves while the case runs.
- A busy machine holds a shorter window than a quiet one, which is why the figure is a property of this machine rather than of macOS.
- Both numbers are one command and belong in the notes at first contact.
This is the record every later absence finding rests on: an absence means nothing until the store is shown to have covered the period.
Read the rows as a collection order rather than a syllabus. The durable stores will wait for you and the top two will not, which is why a triage collection starts at the volatile end regardless of which store the investigation is ultimately about.
That ordering has a cost worth accepting knowingly. Collecting the volatile stores first means spending the early minutes on material that may turn out to be irrelevant, and the alternative is spending them on material that would still have been there tomorrow. Only one of those mistakes is recoverable.
There is a second reason the order runs this way, beyond simple perishability. The volatile stores are the ones that record activity as it happens, so they are also the ones that answer questions about sequence and timing, which is usually what an investigation is actually short of. The durable stores tend to describe state, and state can be inferred later from a well-preserved image in a way that a missing hour of activity cannot. That asymmetry is the whole argument for the ordering, and it holds regardless of which question the investigation eventually turns out to be about.
It is also why the ordering survives being wrong about the case. An examiner who collects the volatile stores first and then discovers the question was about file provenance has lost nothing, because the filesystem records were never going anywhere. The reverse mistake costs the window permanently, and no amount of later diligence recovers it.
The acquisition race
Which the scenario losesPublished practice turns all of this into an operating rule rather than an observation. An investigation that opens days or weeks after a suspected incident may find the relevant entries have already rolled off, and log acquisition is therefore treated as an early-triage priority specifically because of that race.
THE RACE, AS THE SCENARIO RUNS IT
day 0 incident begins, unnoticed
day 11 suspicion raised, plan written assuming 30 days
day 11 plan approved, collection scheduled for next week
day 18 collection runs
what the log has to reach back to cover day 0 18 days
what a busy machine may actually hold under 1 day
The sentence that follows it is the one the scenario ignored: acquiring early preserves the data, and waiting until the formal investigation is scoped often loses the most relevant entries.
So the collection is not a step that happens once the plan is approved. On this platform the collection is what protects the plan from becoming unanswerable, and scheduling it for next week is a decision to investigate with less, taken by somebody who probably did not realize they were taking it.
Raising it is straightforward and rarely resisted once the mechanism is explained. The argument is not that the investigation is urgent in the abstract but that a specific store is discarding entries while the scheduling conversation happens, and that the cost of collecting today and not needing it is a few gigabytes of storage. Framed that way it stops being a request for priority and becomes a statement about what will and will not exist next week. Most people who decline an urgent-sounding request will agree to a cheap one that prevents a known loss. The framing matters because the decision usually sits with somebody who is weighing it against other requests rather than against the evidence.
Having the measurement to hand makes that conversation shorter still. Quoting the actual window on the actual machine, taken minutes earlier, moves the discussion from a general claim about logs to a specific statement about this device, and specific statements are considerably harder to defer than general ones. It also demonstrates that the concern is measured rather than reflexive, which is worth something the next time you raise one.
Measure the window, do not assume it
And it takes one commandWhat to do instead of quoting a figure is to read the window off the machine, and the oldest surviving entry is a fact you can establish rather than estimate. Published research notes that tools exist which report the datestamp at the start of the current collection of persistent log files, which is the indicator of the oldest entry available from them.
# How much space the persistent store is using, and how many files it holds
% sudo du -sh /private/var/db/diagnostics/Persist
% sudo ls -1 /private/var/db/diagnostics/Persist | wc -l
# The oldest entry the store still holds, which is the real window
% log show --style compact --last 30d 2>/dev/null | head -3
Run the first pair before the third, because they tell you whether the folder is at its cap. A folder sitting near the target size is one that has already been discarding, so the oldest entry is the edge of what survives rather than the day the machine was built.
Establishing that window costs a minute and changes what a plan can promise. A window shorter than the incident is not a failure of the collection; it is a finding about scope, and it has to be known before the collection rather than discovered after it, because knowing it early is what lets somebody route the question to a store that still holds an answer.
The measurement is also worth repeating rather than taking once. On a machine still running, the window continues to move while the case does, so a figure established at intake is already out of date by the time a collection is scheduled. Recording both the value and the moment it was taken is what lets anybody reading the plan later work out how much further the edge has traveled since. Two dates and a value, written once, answer a question that otherwise requires going back to the machine.
Going back is often not possible in any case, which is the point. By the time a plan is questioned the machine may have been returned, rebuilt or simply left running for a month, and the measurement you took at intake is the only record that the window was ever a particular size. That makes it evidence about the examination rather than a working note.
Writing the window into the plan
Before somebody assumes itHow the measured window reaches the people making decisions is the last question, and the form matters. State it as a measured value with the date it was measured, because it is shrinking while the case runs and a figure with no measurement date is an assumption wearing a number.
THE RETENTION LINE IN A COLLECTION PLAN
measured the oldest entry actually present, and on what date
mechanism size-capped, so the window moves as the machine runs
covers whether it reaches the suspected start of the incident
if it does not which stores are being relied on instead
a figure with no measurement date is an assumption
Read the block as the retention line a plan should carry. The mechanism line records that the window moves as the machine runs, which is what stops a reader treating the measurement as a constant they can rely on a fortnight later.
THE SAME LIMITATION, TWO WAYS
weak "log retention limited the available period"
strong "oldest surviving entry 06 Mar, measured 14 Mar. Store is size-capped,
so the edge moves daily. Does not reach the suspected 02 Mar start.
Filesystem events and signpost data relied on for that period."
same facts, and only one of them is a plan
Read the two as the same finding written by somebody who measured and by somebody who did not. The second costs four lines and answers the questions the first invites.
The fourth line is the one that turns a limitation into a plan. If the log will not reach, the longer-lived stores and the filesystem records are what the case rests on, and saying so early is what lets somebody collect them properly instead of discovering the gap when the log comes back short.
That framing also changes how the limitation reads to somebody outside the technical work. A plan that names the window, says what it covers and names the fallback stores is describing a considered scope; the same plan without those lines produces a report that appears to have simply missed the period in question. The facts are identical and only one version survives being asked about. Which of the two you produce is decided at the planning stage rather than at the writing stage, which is why this belongs in a section about how the course is ordered rather than in one about reporting.
Everything downstream inherits that ordering. The modules that follow read the stores in roughly the sequence their lifespans dictate, acquisition and logging before filesystem and user activity, and that is a consequence of this section rather than an editorial preference. A course arranged by topic would teach them in a different order and leave students arriving late to the evidence that leaves first.
Practice
Measure your own window hands onAlmost nobody has measured the log window on a machine they own, and the number usually surprises people who have been quoting thirty days.
- Answer the scenario. Name the assumption in the collection plan and say what it costs given an incident eleven days old.
- State the mechanism. Say what the store is capped by, and why that means a busy machine keeps less.
- Order the stores. List them from fastest to slowest and say which end a triage collection starts at.
- Now on a Mac you have permission to use, measure the size of the persistent store and count the files in it.
- Then find your oldest surviving entry. Compare it against the thirty-day figure and write down the difference. That number is the one to carry into the next plan you write.
Published research describes the log store as maintained by total folder size rather than by age, with a target around 520 to 530 MB across roughly fifty files and no apparent way to change that allowance, so the window is a property of how fast a machine fills the space. Published observation records healthy machines once keeping up to twenty days and that period reducing as entry volume rose. Published description states the default rotation is approximately seven days while a busy system can rotate in under twenty-four hours. Published analysis states directly that the thirty-day figure is largely incorrect, with lifetimes ranging from minutes to effectively indefinite, while persistent signpost data lasts weeks to months and activity records run to roughly twenty-eight or thirty days by observation. Published practice treats acquisition as an early-triage priority because waiting often loses the most relevant entries.