Reading width
Wide uses the full column for everything, text, diagrams, code, and exercises. Narrow keeps the standard reading width.
Text size
Scales the body text. Headings and code blocks keep their size.
In this section
Security Data and AI Tools: What Must Never Leave Your Estate
Introduction
Every other section in this module is about whether a generated answer is correct. This one is about the only risk that has nothing to do with correctness: what leaves your organization when you ask for help, and who else can read it afterwards.
The instinct that protects documents and email does not fire on a paste box, because log lines do not look like the kind of thing anyone classifies. They are, and this section shows precisely what an outsider reads from one: your naming convention, your topology, who is attacking you, and sometimes a live weakness. By the end you will know where the line sits, and a fifteen-second redaction that keeps everything the assistant actually needs.
Scenario
An analyst has twelve lines of an unfamiliar log format and forty minutes of vendor documentation ahead of them. They paste the lines into an assistant and ask what they are looking at. Thirty seconds later they have a readable explanation and carry on with the investigation. Nobody was reckless, no policy was consciously broken, and the question of what was actually in those twelve lines never came up, because the lines were the problem rather than the subject.
The paste that starts it
Nobody was reckless and no policy was consciously brokenAn analyst has a log format they do not recognize. Twelve lines, a vendor they have not worked with, and forty minutes of documentation ahead of them. They copy the lines into an assistant and ask what they are looking at.
WHAT WENT INTO THE PASTE BOX
2026-03-02T22:55:12Z|r.scott@ne.com|NE-LEWIS-LT|45.83.64.117|4624|3|Ntlm|SUCCESS
2026-03-02T22:55:09Z|r.scott@ne.com|NE-LEWIS-LT|45.83.64.117|4625|3|Ntlm|50126
2026-03-02T22:54:58Z|svc-sql@ne.com|SRV-NGE-MCR-APP01|10.0.1.44|4624|3|Ntlm|SUCCESS
THE QUESTION WAS "WHAT PRODUCES THIS FORMAT". NONE OF THE
ANSWER DEPENDS ON A SINGLE VALUE ON THESE LINES.
Thirty seconds later they have a readable explanation and they carry on with the investigation.
This is the single most common way security data leaves an organization, and it is done by careful people with good intentions. Nobody is being reckless. The analyst has a problem, there is a tool that solves it, and the question of what was in those twelve lines never quite surfaces because the lines were the problem rather than the subject.
The purpose of this sub is to make that question surface automatically, and to make the answer cheap to act on. Not through a policy nobody reads, but through a habit that costs nothing at the moment it matters.
What is actually in a log line
What an outsider reads from a single recordTake a single row from the Northgate estate and read it as an outsider would.
2026-03-02T22:55:12Z r.scott@ne.com 45.83.64.117 SRV-NGE-MCR-APP01
Office 365 Exchange Online ResultType=0 singleFactorAuthentication
An analyst sees one successful sign-in. Somebody outside your organization sees:
- A real person's work identity, in a format that reveals the naming convention for every other employee
- Your internal hostname convention, which encodes the site, the function and the number, so
SRV-NGE-MCR-APP01tells a reader there is anAPP02and probably aDB01at the same site - Which services you run and which authentication method that service accepted
- The fact that this account authenticated with a single factor, which is a live weakness rather than a historical fact
One line. Now consider that the paste is rarely one line, and that a plausible request is "here are the last two hundred events, what am I looking at".
The Reasonable Mistake
The reason log data is underestimated
Analysts assess disclosure risk by asking whether the data is sensitive, and log lines do not feel sensitive: no salary, no medical record, no customer detail. But the risk here is not that the content is private. It is that the content is a map of your estate: naming conventions, topology, service inventory and current weaknesses, which is precisely the reconnaissance an attacker would otherwise have to work for.
The categories are easier to hold as a picture than as a list, because the only thing that matters is where the line falls between them.
Four separate disclosures in one line, and none of them is the incident the analyst was asking about.
WHAT THEY PASTED WHAT IT TOLD A STRANGER
2026-03-02T22:55:12Z nothing
r.scott@ne.com your naming convention, and a real person
NE-LEWIS-LT your host convention, and one live hostname
45.83.64.117 an address currently attacking you
10.0.1.142 your internal range
4624 3 NTLM SUCCESS nothing you did not want to share
The analyst was not being careless. They were being efficient about a format they did not recognize.
Four categories, and where the line falls
The shape may leave, the specifics may notNot everything is equally sensitive, and a rule that treats it all the same gets ignored on a busy shift. Four categories, in rising order.
1. Generic technical content. Query syntax, a public vendor field name, a CVE identifier, a general question about how a protocol behaves. Nothing here identifies your organization. "How do I express a time window in KQL" discloses that you use KQL, and that is all.
2. Structure without identity. A log format with the values removed, a schema, a rule template. This is where most genuinely useful requests live, and it is almost always safe. "Here is the shape of this log format with the values redacted, which field is likely to hold the source address" gets you the same answer as pasting the real lines.
The Reasonable Mistake
Skipping category 2 on the way to category 3
Almost nobody decides to disclose estate identifiers. What happens is that the redacted version feels like it will not be good enough, so the real lines go in to be safe. The answer to a structural question is identical either way, and the belief that it will not be is the thing doing the damage.
3. Estate identifiers. Usernames, hostnames, internal addresses, application names, group names, file paths containing user directories. This is the category the opening example falls into, and it is the one people misjudge, because individually each item feels trivial.
4. The incident itself. What was compromised, when, what the attacker did, what you have and have not contained. This is the category where disclosure is not merely a privacy question but an operational one: the details of a live incident are the details an attacker would most like to know you know.
Most policies are written as a blanket prohibition, which fails because category 1 and 2 requests are useful, safe and constant, so the prohibition is broken hourly and stops meaning anything.
Drawing the line at estate identifiers instead gives you a rule somebody can follow on a busy shift, because it permits the requests people actually make.
Categories one and two answer nearly every question you would ask an assistant. Three and four are the ones that identify you.
1 GENERIC TECHNICAL syntax, a CVE, how a protocol behaves
costs nothing. It is public already.
2 STRUCTURE, REDACTED a format with USER1 and HOST1 in it
costs nothing, and it answers the question
------------------ THE LINE IS HERE ------------------
3 ESTATE IDENTIFIERS real accounts, hostnames, ranges
costs your naming convention and topology,
permanently, from one paste
4 THE INCIDENT ITSELF what is compromised, and what you know
costs the thing an attacker most wants to
know: whether you have noticed
That is the whole classification, and it is short enough to run in your head at the moment you are about to paste something.
The questions that decide it
What to ask before pasting anythingThree questions, in order, answerable in about ten seconds.
Does this identify us? Not "is it secret" but "could a reader work out which organization this is and how it is built". Hostnames and internal addresses answer yes.
Would this help someone attacking us? A naming convention helps. A list of applications helps. Knowing which account authenticates without MFA helps a great deal.
Can I get the same answer without it? Nearly always yes, and this is the question that resolves the majority of cases. The analyst in the opening example did not need the values. They needed to know what the format was. A single line with the values replaced by placeholders gets an identical answer.
Redaction is faster than the deliberation. The realistic objection is that redacting takes time you do not have. It takes about fifteen seconds, which is less than the time spent wondering whether the paste is allowed.
Test it on your own assistant
and what does each field mean?
TIMESTAMP|USER1|HOST1|10.0.0.1|4624|3|NTLM|SUCCESS
What this demonstrates. The structural work an assistant does on an unfamiliar format almost never depends on the values. Substituting USER1 and HOST1 costs you nothing and takes fifteen seconds, and it is the difference between sending a format and sending a map of your estate.
Try it a second time with the values restored and compare the answers. Here are the two lines side by side.
WHAT YOU HAVE 2026-03-02T22:55:12Z r.scott@ne.com 45.83.64.117 4624 3 NTLM SUCCESS
WHAT YOU SEND TIMESTAMP USER1 HOST1 10.0.0.1 4624 3 NTLM SUCCESS
WHAT SURVIVES the event ID, the logon type, the auth package, the outcome
WHAT LEAVES nothing that identifies you
Every token the answer actually used is in the second line. If the two answers match, and they usually do, you have measured the cost of the safe path and it is zero.
Making the safe path the fast path
Why redaction does not cost you the answerRules that depend on judgment under pressure fail under pressure, which is when they matter.
Have an approved tool. The single largest driver of data going somewhere it should not is the absence of an approved place for it to go. An analyst with a sanctioned assistant that meets the organization's data terms does not paste into a consumer service; an analyst with nothing sanctioned finds their own.
Answer the question before it is asked. "May I use this for log analysis" needs a documented answer that an analyst can find in under a minute at two in the morning. An unanswered question resolves itself in favor of the shortcut.
THE QUESTION AN ANALYST ASKS AT 02:40
"can I paste this into the assistant?"
answer findable in under a minute -> they follow it
answer requires asking someone -> they decide alone
answer does not exist -> they decide alone
Two of the three outcomes are identical, and neither is policy.
Default to structure. Build the habit of asking about the shape rather than the content, because it is faster to redact by reflex than to assess each case. The reflex also survives fatigue, which the assessment does not.
Each control removes the moment of choice rather than asking somebody to choose well at two in the morning.
Assume permanence. Whether a given service retains prompts, trains on them, or logs them for abuse monitoring is a question with a different answer per vendor, per tier and per year. The stable working assumption is that anything sent may persist somewhere you cannot reach, and that assumption costs you nothing if it turns out to be pessimistic.
Before pasting anything, in this order:
1 IS THERE AN APPROVED TOOL? if not, that is the finding,
not this paste
2 IS THE ANSWER FINDABLE? if you cannot check the policy
in a minute, assume the strict
reading
3 CAN I ASK ABOUT SHAPE? replace every identifier with
USER1, HOST1, 10.0.0.1
4 WOULD I MIND THIS PERSISTING? assume it does. If that changes
the answer, do not send it.
Redaction is a substitution, not a deletion. Replace consistently rather than removing: USER1, USER2, HOST1, 10.0.0.1. Consistency matters more than obscurity, because the relationships between the values are often the point, and "the same user appears in lines 1, 4 and 9" survives a consistent substitution perfectly well. Section 07 covers why the substitution costs you nothing.
What a good sanctioned tool actually needs
Four operational questions, and the one that is not on the listIf your organization is choosing what to approve, the questions that matter operationally are narrower than a procurement questionnaire suggests.
Does it train on your inputs by default, and can that be turned off in writing? The distinction between a consumer tier and a business tier is usually exactly this, and it is usually the whole of the difference that matters.
SAME LINE, TWICE
sent as-is 2026-03-02T22:55:12Z r.scott@ne.com NE-LEWIS-LT 4624 3 NTLM SUCCESS
sent redacted TIMESTAMP USER1 HOST1 4624 3 NTLM SUCCESS
the answer "Windows security event 4624, logon type 3 is a network logon,
NTLM authentication, successful. This is a remote session
against the host using NTLM rather than Kerberos."
identical from both. Every token the answer used is in the second line.
What is retained, for how long, and who can read it? Abuse monitoring is legitimate and it means a human at the vendor can potentially read a prompt. That is acceptable for category 1 and 2 content and is a genuine consideration for category 4.
RETENTION how long are prompts kept, and by whom TRAINING are prompts used to train, and can that be off ACCESS can a human at the vendor read a prompt, and when JURISDICTION where does the data sit, and does that matter here TIER does the answer change on the paid plan Every one of these has a different answer per vendor, per tier and per year, which is why it is a document rather than a memory. The working assumption when it is unanswered: anything sent may persist somewhere you cannot reach.
Where does it run? Data residency is a real constraint for some organizations and irrelevant for others, and it is worth establishing which you are before the conversation rather than during it.
THE ANSWER YOU NEED, IN THE FORM AN ANALYST CAN USE
SANCTIONED [tool name], enterprise tier, no training on prompts
FOR log formats, query help, technique explanation
NOT FOR real hostnames, accounts, addresses, incident detail
IF UNSURE redact and send. If redaction breaks the question,
ask a person instead.
WHERE pinned in the SOC channel, and in the runbook
Five lines. Findable in twenty seconds at two in the morning.
Can you demonstrate the terms to an auditor? Not whether you believe the terms are adequate, but whether you can produce them. The question is asked after an incident and it is asked in writing.
Conflating them is how organizations end up with a tool that is safe and useless, which analysts then route around.
If it has already happened
What to do, and what not to bother doingIt has. Not necessarily to you, and if you manage a team, to somebody on it. Treat that as the starting condition rather than the failure state.
The useful response is not disciplinary. It is to establish what category was disclosed, because that determines whether anything needs doing at all. Category 1 or 2 needs nothing. Category 3 is worth recording, because a naming convention disclosed once is disclosed permanently and it belongs in your threat model rather than in a disciplinary file. Category 4, during a live incident, is a decision for whoever is running the incident and probably for legal.
The three on the left leave the cause in place. Only the right-hand column changes what happens the next time somebody is stuck.
The failure mode to avoid is silence. An analyst who believes disclosure ends their career does not report it, which converts a small, containable event into an unknown one. Whatever the policy says, the operational goal is that people tell you, and that goal is set by how the first report is handled rather than by what is written down.
1 ESTABLISH THE CATEGORY 1 or 2, nothing to do
3, record it, it is now permanent
4, whoever runs the incident decides
2 ROTATE ANYTHING LIVE credentials, tokens, keys in the paste
3 SANCTION A TOOL the absence of one is the finding
DO NOT run a disciplinary process
ask people to be more careful
try to retract the prompt
ALL THREE LEAVE THE CAUSE IN PLACE, AND THE THIRD DOES NOT WORK.
Module 7 develops this into a working policy: what to sanction, how to write terms that survive a vendor change, and how to build a workflow where the compliant route is genuinely the quickest one available.
Why redaction does not cost you the answer
What the assistant actually reads, and why the substitution is nearly freeThe objection to redacting is that it removes the thing the assistant needed, and understanding why it usually does not is what makes the habit stick.
A model reading an unfamiliar log line is doing structural work. This token is positional, that one is a timestamp, the third has the shape of an identifier, the delimiters suggest a particular vendor's format. Almost none of that depends on the values themselves. Replace every username with USER1 and the structural answer comes back identical, because the question was about shape.
The tokens that identify you and the tokens that answer the question are different tokens, which is why the substitution is nearly free.
Read the left column and ask which of those tokens you could not have supplied from a vendor manual.
DO IT ON THE LINE FROM THE OPENING PASTE. FIFTEEN SECONDS.
BEFORE 2026-03-02T22:55:12Z r.scott@ne.com NE-LEWIS-LT 45.83.64.117 4624 3 Ntlm SUCCESS
AFTER TIMESTAMP USER1 HOST1 10.0.0.1 4624 3 Ntlm SUCCESS
STILL THERE the event ID, the logon type, the auth package,
the outcome. Every token the answer is built from.
GONE who you are, and how you name things.
Where values do matter, they are the ones that are safe to keep. A result code, an event ID, a protocol name, an error string: these are vendor vocabulary rather than estate detail, and they are precisely the tokens the assistant needs to recognize the format.
Replace consistently rather than deleting. The relationships between values are often the point: "the same user appears in lines 1, 4 and 9" survives redaction perfectly well if that user is USER1 throughout, and it is frequently the fact that mattered. Deleting the values destroys that; substituting them does not.
Practice
Redact one real log line, and time it hands onReading this section changes nothing. Fifteen seconds of redaction on one real line is what turns it into a reflex, and the reflex is the only part that survives a busy shift.
- Take a log line you would genuinely have pasted somewhere.
- Replace consistently: usernames to
USER1, hostnames toHOST1, addresses to10.0.0.1. - Keep result codes, event IDs and error strings. Those are vendor vocabulary and the assistant needs them.
Then find out what your organization has actually sanctioned, and whether the answer is written down anywhere an analyst could find it at two in the morning. An unanswered question resolves itself in favor of the shortcut, every time.
Next: section 0.7 is the drill. Four generated queries against live data, three of them wrong and one of them fine.