Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
The Job This Course Is For
Introduction
Endpoint security splits into two jobs that are bought together and done by different people. One builds the estate: onboarding, policy, hardening, the controls and their exceptions. The other operates it: what fires, what it means, what you do about it, and what an investigation will still be able to establish next month.
This course is the second job. It assumes the first is done, and it is written for the practitioner who arrived after the deployment and is now answerable for what the estate detects.
You will finish this section able to say which of the two jobs a given piece of work belongs to, which matters more than it sounds: most of the friction between a security team and a platform team is one of them being asked to do the other's job without the access to do it.
Scenario
You inherit an estate with 865 endpoints, twelve servers and a Defender deployment somebody finished eight months ago. The console is green. In your first week an alert fires that nobody can explain, a hunt you run returns four thousand rows, and a request arrives for evidence from a machine that was rebuilt on Tuesday. None of those three is a deployment problem, and none of them is fixed by changing a policy.
Two Jobs, One Console
And why the split is not about difficultyThe obvious way to divide endpoint security is by difficulty, with the basics in one course and the advanced material in another. That division produces two halves that each teach part of a discipline and neither of which is a job anybody holds.
The division that matches reality is engineering against operations. Engineering decides what the estate does; operations decides what the organization does about what the estate reports. They use the same console and almost none of the same skills.
Figure EO0.1a. Almost everything difficult in the operations job traces back to the arrow: settings inherited without the reasoning that produced them.
The two boxes are staffed differently in most organizations and share one console, which is where the confusion starts.
Microsoft Defender portal
The page both jobs open, and it is worth noticing what each of them reads on it. The engineering job reads onboarding status and health; the operations job reads last-seen, device group and criticality. Same page, two different columns, and neither team usually knows which columns the other one uses.
Both teams are reading the same rows and drawing different conclusions from them, which is the arrow in that figure made concrete.
Reading Configuration as Evidence
What a setting tells you about a decision nobody wrote downThe arrow is where this course starts. An operator who can read the estate's current configuration can work out what was decided; an operator who cannot is guessing about the meaning of everything the console reports.
That is worth sitting with, because it is the difference between two operators looking at the same alert. One knows that automatic remediation is set to semi-automatic on this device group, so the alert in front of them means a file was quarantined and a person has to approve the rest. The other does not, and reads the same alert as a thing that has already been handled. Neither of them changed a setting; one of them can act and the other is guessing.
Microsoft Defender portal
With Settings › Endpoints › Advanced features, these two pages hold most of what the handover did not mention: which groups exist and in what order, what automation level each runs at, and which capabilities are switched on for the tenant. Read them in your first week, before you need any of it.
The settings that matter most to an operator are almost never the ones that get documented at handover. Retention windows, automation levels, which device groups exist and in what order, what is excluded and why: these are the inputs to every decision this course teaches, and they live in a console rather than in a runbook.
None of that is a criticism of whoever ran the deployment. A deployment is judged on coverage and on not breaking anything, and both of those are satisfied without ever writing down why a particular exclusion was granted. The reasoning existed; it was in somebody's head during a meeting, and the artifact that survived is the setting.
Which means the first operational skill is reading configuration as evidence of a decision. An exclusion covering a whole volume was granted for a reason, and the reason is either still true or it is a finding. The setting cannot tell you which, but it tells you where to ask.
Two kinds of setting are worth separating as you read. Some were chosen, and somebody weighed something to arrive at them. Others are defaults nobody has ever looked at, and a default is not a decision even though it has exactly the same effect on the machine. Distinguishing them is usually possible: a value that differs from the product default was chosen, and a value that matches it probably was not.
Treat that as the posture to arrive with rather than as a task. An operator who reads the estate as a set of decisions asks better questions of the people who made them, and gets better answers, because a question about a specific exclusion granted in March is answerable where a general request for documentation is not.
It also changes what a handover conversation is for. The useful version is not a walkthrough of the console, which you can do yourself, but a list of the decisions somebody remembers making and the reasons they remember. Those reasons decay fast: people leave, and a setting outlives the person who set it by years.
Where nobody is left to ask, the configuration is still evidence. An exclusion path naming a product tells you which team asked. A device group created before the others tells you the order things were built in. Reading an estate this way is a skill in its own right, and it is most of what a first month should produce.
The Work Is Reactive and the Preparation Is Not
Which is why half this course happens before anything firesOperations sounds reactive, and the visible part of it is. An alert fires, somebody triages it, a decision gets made. That part is unpredictable and it is what the job looks like from outside.
The part that decides how well it goes is entirely predictable and happens beforehand. Whether the telemetry that answers a question was being collected. Whether the retention window still covers the period in question. Whether a rule was written for the technique or for one tool's implementation of it. Whether anybody can pull an artifact off a machine at speed, and whether the machine still exists.
Figure EO0.1b. The analyst's skill is real and it operates on whatever the right-hand column left them.
Read the right-hand column. The analyst's skill is real and it operates on whatever the left-hand column left them, which is why a course on endpoint operations that only taught triage would be teaching the smaller half.
Three of those five have no owner in most organizations, which is the more useful observation. Retention belongs to whoever holds the license budget. Script block logging belongs to whoever writes group policy. The live response role belongs to whoever administers the security portal. None of them belongs to the person the alert wakes up, and all three decide what that person can do.
There is a version of this conversation that works and a version that does not. Asking for better logging in general gets nowhere, because it is a cost with no attached consequence. Naming the specific alert whose investigation stalled, and the specific setting that would have prevented the stall, is a different request: it has a case behind it, and the case is one somebody already had to explain.
Microsoft Defender portal
Shows who holds which response capability, and Permissions › Roles shows the same for the wider portal. Neither is a page an operator is usually pointed at, and both answer a question that otherwise gets discovered during an incident: whether the person on call tonight can act.
This course puts them in front of you for that reason. You may not own the settings, and you can always establish what they currently are, which turns an unanswerable question during an incident into a known constraint before one.
The distinction between owning a setting and knowing it is worth being precise about, because it decides what you can promise. An operator who knows retention is thirty days can tell an investigator on the first call which questions are answerable, which is enormously more useful than finding out on day four. That is not authority over the setting; it is knowledge of it, and it costs one query.
Why the Console Is Not the Job
Three questions it does not answerAn endpoint security console is very good at telling you what it found. It is structurally unable to tell you what it did not look for, and most of the operations job lives in that gap.
Three questions recur through every module in this course, and none of them has a console view.
Figure EO0.1b. The left column is what an operator is actually asked in a review. The right column is where this course answers it.
Read it knowing that limit. It tells you about absence within the inventory and nothing about absence from it, which are two different problems with two different owners.
Absence is also harder to raise with anybody else. A queue of two hundred alerts is visible to a manager and reads as a workload; an estate collecting nothing from forty machines is invisible to everybody and reads as quiet. Making absence legible to people who do not run the console is a large part of what an endpoint operator is for, and it is the reason several sections of this course end with a written artifact rather than a query.
The distinction that matters
A console reports presence. It is very good at telling you what it found, and everything it shows you is something that happened and was recorded.
Every question in that figure is about absence. What is not collected, what a rule did not catch, what an investigation will not be able to establish. No view reports those, because a view is built from records and absence has none.
The written artifact is the point rather than a formality. A number in a document survives the person who found it and can be re-run next quarter; the same number said in a meeting is gone by Friday. Several of the most useful outputs of this course are one page long and get read by somebody who will never open the console.
Absence Is the Recurring Shape
And it is harder to raise than a queue that is too longNotice that all three are questions about absence. That is the recurring shape of this discipline: the alert queue shows you presence, and almost everything that goes wrong in endpoint operations is something that was not there.
Microsoft Defender portal
The closest the console comes to showing absence: sensor health, onboarding status, and devices that have stopped reporting. Read it knowing its limit, which is that it can only describe devices the platform already knows about.
Absence within the inventory and absence from the inventory are two different problems with two different owners, and only the first of them has a page.
There is a practical consequence for how you spend a first month on an unfamiliar estate. The instinct is to work the queue, because the queue is in front of you and it is measurable. The queue tells you about the detections that exist and fire; it is silent about everything else, so a month of good queue work can leave you knowing less about the estate than an afternoon spent reading what it collects.
Do both, in that order. Read the collection position first and the queue second, because the queue only means something once you know what it is drawn from.
None of that means the queue is unimportant. It is where the job is judged and it is where the organization notices whether you exist. The point is narrower: the queue is an output of the collection position, so it is not the place to learn what the collection position is.
There is one exception worth knowing about. If an incident is running, work the incident. The reading-first advice is about the ordinary weeks, and an estate in the middle of a compromise is not one of them.
The same ordering runs through the whole course, which is why the modules sit in the order they do. Telemetry before hunting, because a hunt is a question asked of data that either exists or does not. Hunting before detection engineering, because a detection is a hunt somebody decided to run forever. Readiness before triage, because collection decided during an incident is collection you did not get.
That last pair was the other way round until recently and it produced a predictable result: students learned to investigate well and then learned, afterwards, which of the things they had just been taught to look for were not being recorded on their own estate. The order now matches the order the decisions have to be taken in, which is not always the order they feel natural in.
What the Operations Job Produces
Four outputs, and who consumes eachThe job has four outputs and it is worth being concrete about them, because "monitoring the estate" describes none of them and is what the role is usually called.
Figure EO0.1c. Read the right-hand column rather than the left. Each of the four fails in a specific and recognizable way.
Read the failure column rather than the output column. Each of the four fails in a way you can recognize on sight, and the six modules after this one are arranged around preventing those four rather than around the tools that produce the outputs.
The distinction that matters
Three of the four are produced under time pressure, by somebody with an incident in front of them. They get done because the situation demands them, badly or well.
The fourth is produced before anything has happened, by somebody with nothing in front of them. Nothing demands it, nobody is waiting for it, and it is the only one of the four with no second chance once it is skipped.
Note that three of the four are produced under time pressure and one is not. The preserved artifact is the odd one out: it is produced before anything has happened, by somebody with no incident in front of them, which is exactly why it is the one most often skipped.
Two Failures That Get Misdiagnosed
The hunt with the wrong question, and the correct decision recorded badlyTwo of those failures are worth expanding now, because they are the two that get misdiagnosed most often.
A hunt that returns four thousand rows is usually called a tuning problem and treated by adding filters until the number is small. That is sometimes right and it is frequently how a hunt gets narrowed until it only returns things you already knew about. The other reading is that the question was wrong: a hunt asks what is unusual here, and four thousand rows means the thing you picked is not unusual, so the filter you need is a better question rather than more conditions.
The distinction that matters
A tuning problem is a volume you can define your way out of. You know what a suspicious row looks like, there are simply too many rows, and filters get you to the ones you meant.
A question problem is a volume no filter reaches. You would recognize the answer on sight and cannot describe it in advance, so every filter you add removes rows at random with respect to the thing you are looking for.
The test that separates the two is whether you can say, in advance, what a suspicious row would look like. If you can, the volume is a tuning problem and filters are the answer. If you cannot, no amount of filtering will help, because you will recognize the answer only when you see it and four thousand rows is too many to see.
A triage decision with no evidence behind it is rarely a careless analyst. It is almost always an analyst who reached a correct conclusion from something they could see at the time and cannot now reproduce, because the view they used has moved on. The fix is a record made during the decision rather than more rigor afterwards, and that is a habit rather than a skill.
The fourth failure, a machine rebuilt before anything was collected from it, is the one nobody in security controls. Rebuilding a compromised machine quickly is correct behavior from a service desk measured on getting a user working again, and it destroys the only copy of the evidence. That is not solved by asking them to stop; it is solved by collection that happens automatically on a signal they already generate, which is what makes readiness an engineering-adjacent activity done by operations.
It is worth understanding why that one sits in the middle of this course rather than at the end. Readiness is the only one of the four failures that cannot be fixed after it happens: a bad detection can be rewritten and a poor hunt re-run, and an artifact that was never collected from a machine that no longer exists is simply gone. Everything else in the job has a second chance and that one does not.
Who This Is Written For
And who should be reading the other courseThis is written for somebody responsible for what an endpoint estate detects and how the organization responds to it. A SOC analyst moving from triaging alerts to writing the rules that produce them. A detection engineer who has worked identity or network telemetry and now owns endpoints. An incident responder who wants the estate ready before the next call rather than after it.
It assumes an estate that is already deployed, onboarded and reporting. If yours is not, the sections here will read as though they are describing somebody else's environment, because they are: every reading in this course queries telemetry that a deployment produces, and there is no useful way to query telemetry that does not exist yet.
That is not a gate on your experience. There is no minimum, and every concept is introduced at first use. It is a statement about the estate rather than about you, and the distinction matters because the fix is a deployment rather than a prerequisite course.
What makes it go faster is access rather than experience. Advanced hunting, the ability to read a device timeline and sight of the configuration are the three things that turn a section of this course from a description into a reading you take on your own estate. None is required, and every reading in the course is shown with its output.
It is also worth saying what this course does not cover, because two adjacent disciplines look like they belong here and do not. Building the estate is ARC402, and that includes onboarding, policy, hardening and the exception register. Deep host forensics, disk and memory analysis at the artifact level, is a forensics course; this one covers preserving what such an investigation would need, which is a different job done at a different time by frequently a different person.
One boundary is deliberately soft. The telemetry module immediately after this one restates the sensor and data model that the engineering course also covers, briefly and from the operator's side. That is duplication and it is intentional: a student holding only this course must be able to understand the data they are querying without buying another one, and a course that sent them elsewhere for the meaning of its own tables would be incomplete rather than efficient.
- Does it change what the estate does? Onboarding, policy, hardening and exceptions are engineering, and they change the machine.
- Does it change what the organization does about what the estate reports? Detection, hunting, triage and preservation are operations, and they change a decision.
- Does it need access you do not have? That is the most reliable signal that the work belongs to the other side of the handover.
- Is somebody asking you to do both? Common, workable, and worth naming out loud so the time is budgeted for two jobs rather than one.
One caveat on the test, because it is not a way to refuse work. Plenty of operators do engineering tasks, and on a small team the same person does both jobs on the same afternoon. The value of naming which is which is in the estimate rather than in the refusal: engineering work needs a change window, a test and somebody to approve it, and operations work usually does not. A week planned as though both were the same kind of work will overrun on the half that needs the change process.
Run that test on the scenario at the top of this section and all three items land on the operations side, which is why none of them was fixed by changing a policy. The unexplained alert is a detection with no documented scope. The four thousand rows are a hunt with the wrong question. The machine rebuilt on Tuesday is a readiness decision nobody took. Three modules of this course, one week of somebody's job.
Practice
Sort your own week hands onEverything above describes the job. This puts a week of your own against it.
- List the endpoint security work you did last week, at the level of individual tasks rather than projects.
- Sort each into engineering or operations using the test above, and mark the ones you could not complete because of access.
- Count the four outputs. How many detection rules, hunts, triage decisions and preserved artifacts did that week actually produce?
- Name the failure you saw most. One of the four failure modes above will be more familiar than the others, and it tells you which module to read first if you do not read in order.
Nothing in that week required a tool the estate did not have. Each of the three needed a decision taken earlier by somebody who knew it was a decision, and in each case the decision was cheap at the time and expensive once the week had started.
The next section takes the assumption this one just made, that the estate is already deployed, and says exactly what that means in practice.