In this section

Security Data and AI Tools: What Must Never Leave Your Estate

Module 0

Introduction

Every other section in this module is about whether a generated answer is correct. This one is about the only risk that has nothing to do with correctness: what leaves your organization when you ask for help, and who else can read it afterwards.

The instinct that protects documents and email does not fire on a paste box, because log lines do not look like the kind of thing anyone classifies. They are, and this section shows precisely what an outsider reads from one: your naming convention, your topology, who is attacking you, and sometimes a live weakness. By the end you will know where the line sits, and a fifteen-second redaction that keeps everything the assistant actually needs.

Scenario

An analyst has twelve lines of an unfamiliar log format and forty minutes of vendor documentation ahead of them. They paste the lines into an assistant and ask what they are looking at. Thirty seconds later they have a readable explanation and carry on with the investigation. Nobody was reckless, no policy was consciously broken, and the question of what was actually in those twelve lines never came up, because the lines were the problem rather than the subject.

01

The paste that starts it

Nobody was reckless and no policy was consciously broken

An analyst has a log format they do not recognize. Twelve lines, a vendor they have not worked with, and forty minutes of documentation ahead of them. They copy the lines into an assistant and ask what they are looking at.

WHAT WENT INTO THE PASTE BOX
 
2026-03-02T22:55:12Z|r.scott@ne.com|NE-LEWIS-LT|45.83.64.117|4624|3|Ntlm|SUCCESS
2026-03-02T22:55:09Z|r.scott@ne.com|NE-LEWIS-LT|45.83.64.117|4625|3|Ntlm|50126
2026-03-02T22:54:58Z|svc-sql@ne.com|SRV-NGE-MCR-APP01|10.0.1.44|4624|3|Ntlm|SUCCESS
 
THE QUESTION WAS "WHAT PRODUCES THIS FORMAT". NONE OF THE
ANSWER DEPENDS ON A SINGLE VALUE ON THESE LINES.

Thirty seconds later they have a readable explanation and they carry on with the investigation.

This is the single most common way security data leaves an organization, and it is done by careful people with good intentions. Nobody is being reckless. The analyst has a problem, there is a tool that solves it, and the question of what was in those twelve lines never quite surfaces because the lines were the problem rather than the subject.

The purpose of this sub is to make that question surface automatically, and to make the answer cheap to act on. Not through a policy nobody reads, but through a habit that costs nothing at the moment it matters.

02

What is actually in a log line

What an outsider reads from a single record

Take a single row from the Northgate estate and read it as an outsider would.

2026-03-02T22:55:12Z  r.scott@ne.com  45.83.64.117  SRV-NGE-MCR-APP01
Office 365 Exchange Online  ResultType=0  singleFactorAuthentication

An analyst sees one successful sign-in. Somebody outside your organization sees:

  • A real person's work identity, in a format that reveals the naming convention for every other employee
  • Your internal hostname convention, which encodes the site, the function and the number, so SRV-NGE-MCR-APP01 tells a reader there is an APP02 and probably a DB01 at the same site
  • Which services you run and which authentication method that service accepted
  • The fact that this account authenticated with a single factor, which is a live weakness rather than a historical fact

One line. Now consider that the paste is rarely one line, and that a plausible request is "here are the last two hundred events, what am I looking at".

The Reasonable Mistake

The reason log data is underestimated

Analysts assess disclosure risk by asking whether the data is sensitive, and log lines do not feel sensitive: no salary, no medical record, no customer detail. But the risk here is not that the content is private. It is that the content is a map of your estate: naming conventions, topology, service inventory and current weaknesses, which is precisely the reconnaissance an attacker would otherwise have to work for.

The categories are easier to hold as a picture than as a list, because the only thing that matters is where the line falls between them.

ONE LINE. WHAT AN OUTSIDER READS FROM IT. r.scott@ne.com 45.83.64.117 SRV-NGE-MCR-APP01 singleFactorAuthentication Naming convention first-initial dot surname. Now they have every employee. Who is attacking you and that you have noticed Topology site, function, index. There is an APP02 and probably a DB01. A live weakness this account can authenticate without a second factor The risk is not that the content is private. It is that the content is a MAP, and a paste is rarely one line.

Four separate disclosures in one line, and none of them is the incident the analyst was asking about.

WHAT THEY PASTED                    WHAT IT TOLD A STRANGER
 
2026-03-02T22:55:12Z                nothing
r.scott@ne.com                      your naming convention, and a real person
NE-LEWIS-LT                         your host convention, and one live hostname
45.83.64.117                        an address currently attacking you
10.0.1.142                          your internal range
4624 3 NTLM SUCCESS                 nothing you did not want to share
Six tokens on one line. Two are the answer to the question and four are a description of your estate.

The analyst was not being careless. They were being efficient about a format they did not recognize.

03

Four categories, and where the line falls

The shape may leave, the specifics may not

Not everything is equally sensitive, and a rule that treats it all the same gets ignored on a busy shift. Four categories, in rising order.

1. Generic technical content. Query syntax, a public vendor field name, a CVE identifier, a general question about how a protocol behaves. Nothing here identifies your organization. "How do I express a time window in KQL" discloses that you use KQL, and that is all.

2. Structure without identity. A log format with the values removed, a schema, a rule template. This is where most genuinely useful requests live, and it is almost always safe. "Here is the shape of this log format with the values redacted, which field is likely to hold the source address" gets you the same answer as pasting the real lines.

The Reasonable Mistake

Skipping category 2 on the way to category 3

Almost nobody decides to disclose estate identifiers. What happens is that the redacted version feels like it will not be good enough, so the real lines go in to be safe. The answer to a structural question is identical either way, and the belief that it will not be is the thing doing the damage.

3. Estate identifiers. Usernames, hostnames, internal addresses, application names, group names, file paths containing user directories. This is the category the opening example falls into, and it is the one people misjudge, because individually each item feels trivial.

4. The incident itself. What was compromised, when, what the attacker did, what you have and have not contained. This is the category where disclosure is not merely a privacy question but an operational one: the details of a live incident are the details an attacker would most like to know you know.

Most policies are written as a blanket prohibition, which fails because category 1 and 2 requests are useful, safe and constant, so the prohibition is broken hourly and stops meaning anything.

Drawing the line at estate identifiers instead gives you a rule somebody can follow on a busy shift, because it permits the requests people actually make.

WHERE THE LINE ACTUALLY GOES 1 Generic technical content Syntax, a CVE, how a protocol behaves 2 Structure without identity A format with the values redacted 3 Estate identifiers Usernames, hostnames, internal addresses 4 The incident itself What is compromised, and what you know The shape may leave. The specifics may not.

Categories one and two answer nearly every question you would ask an assistant. Three and four are the ones that identify you.

Keep this The four categories, and what each one costs to disclose
1  GENERIC TECHNICAL      syntax, a CVE, how a protocol behaves
   costs                  nothing. It is public already.
 
2  STRUCTURE, REDACTED    a format with USER1 and HOST1 in it
   costs                  nothing, and it answers the question
 
   ------------------ THE LINE IS HERE ------------------
 
3  ESTATE IDENTIFIERS     real accounts, hostnames, ranges
   costs                  your naming convention and topology,
                          permanently, from one paste
 
4  THE INCIDENT ITSELF    what is compromised, and what you know
   costs                  the thing an attacker most wants to
                          know: whether you have noticed
The line does not fall where a sensitivity label would put it. It falls between the shape of a record and the values inside it.

That is the whole classification, and it is short enough to run in your head at the moment you are about to paste something.

04

The questions that decide it

What to ask before pasting anything

Three questions, in order, answerable in about ten seconds.

Does this identify us? Not "is it secret" but "could a reader work out which organization this is and how it is built". Hostnames and internal addresses answer yes.

Would this help someone attacking us? A naming convention helps. A list of applications helps. Knowing which account authenticates without MFA helps a great deal.

Can I get the same answer without it? Nearly always yes, and this is the question that resolves the majority of cases. The analyst in the opening example did not need the values. They needed to know what the format was. A single line with the values replaced by placeholders gets an identical answer.

Question three resolves the majority of cases, which is why redaction is the habit worth building rather than the assessment.

Redaction is faster than the deliberation. The realistic objection is that redacting takes time you do not have. It takes about fifteen seconds, which is less than the time spent wondering whether the paste is allowed.

Test it on your own assistant

Try this Watch how little it needs
Here is a log line format I do not recognize. What system produces it,

and what does each field mean?

TIMESTAMP|USER1|HOST1|10.0.0.1|4624|3|NTLM|SUCCESS

What to look at. That the answer is complete and useful. Every identifier is redacted and nothing was lost: the event ID, the logon type, the authentication package and the outcome are all vendor vocabulary, and they are what the answer was built from.

What this demonstrates. The structural work an assistant does on an unfamiliar format almost never depends on the values. Substituting USER1 and HOST1 costs you nothing and takes fifteen seconds, and it is the difference between sending a format and sending a map of your estate.

Try it a second time with the values restored and compare the answers. Here are the two lines side by side.

WHAT YOU HAVE     2026-03-02T22:55:12Z r.scott@ne.com 45.83.64.117 4624 3 NTLM SUCCESS
 
WHAT YOU SEND     TIMESTAMP USER1 HOST1 10.0.0.1 4624 3 NTLM SUCCESS
 
WHAT SURVIVES     the event ID, the logon type, the auth package, the outcome
WHAT LEAVES       nothing that identifies you

Every token the answer actually used is in the second line. If the two answers match, and they usually do, you have measured the cost of the safe path and it is zero.

05

Making the safe path the fast path

Why redaction does not cost you the answer

Rules that depend on judgment under pressure fail under pressure, which is when they matter.

The four controls below work because none of them asks anybody to resist a shortcut at the moment they are under pressure.

Have an approved tool. The single largest driver of data going somewhere it should not is the absence of an approved place for it to go. An analyst with a sanctioned assistant that meets the organization's data terms does not paste into a consumer service; an analyst with nothing sanctioned finds their own.

Answer the question before it is asked. "May I use this for log analysis" needs a documented answer that an analyst can find in under a minute at two in the morning. An unanswered question resolves itself in favor of the shortcut.

THE QUESTION AN ANALYST ASKS AT 02:40
 
  "can I paste this into the assistant?"
 
  answer findable in under a minute    -> they follow it
  answer requires asking someone       -> they decide alone
  answer does not exist                -> they decide alone
 
Two of the three outcomes are identical, and neither is policy.

Default to structure. Build the habit of asking about the shape rather than the content, because it is faster to redact by reflex than to assess each case. The reflex also survives fatigue, which the assessment does not.

RULES THAT NEED WILLPOWER FAIL WHEN IT MATTERS AN APPROVED TOOL nothing sanctioned means own choice A FINDABLE ANSWER under a minute, at two in the morning REDACT BY REFLEX a reflex survives fatigue. judgment does not ASSUME PERMANENCE costs nothing if it turns out pessimistic None of these asks anybody to resist a shortcut at the moment they are under pressure. The unanswered question is the one that resolves itself in favor of the shortcut.

Each control removes the moment of choice rather than asking somebody to choose well at two in the morning.

Assume permanence. Whether a given service retains prompts, trains on them, or logs them for abuse monitoring is a question with a different answer per vendor, per tier and per year. The stable working assumption is that anything sent may persist somewhere you cannot reach, and that assumption costs you nothing if it turns out to be pessimistic.

Reference The four questions, in order ~15 seconds
Before pasting anything, in this order:
 
1  IS THERE AN APPROVED TOOL?     if not, that is the finding,
                                  not this paste
 
2  IS THE ANSWER FINDABLE?        if you cannot check the policy
                                  in a minute, assume the strict
                                  reading
 
3  CAN I ASK ABOUT SHAPE?         replace every identifier with
                                  USER1, HOST1, 10.0.0.1
 
4  WOULD I MIND THIS PERSISTING?  assume it does. If that changes
                                  the answer, do not send it.
Question 3 answers question 4 most of the time, which is why redaction is the habit worth building rather than the assessment.

Redaction is a substitution, not a deletion. Replace consistently rather than removing: USER1, USER2, HOST1, 10.0.0.1. Consistency matters more than obscurity, because the relationships between the values are often the point, and "the same user appears in lines 1, 4 and 9" survives a consistent substitution perfectly well. Section 07 covers why the substitution costs you nothing.

06

What a good sanctioned tool actually needs

Four operational questions, and the one that is not on the list

If your organization is choosing what to approve, the questions that matter operationally are narrower than a procurement questionnaire suggests.

Does it train on your inputs by default, and can that be turned off in writing? The distinction between a consumer tier and a business tier is usually exactly this, and it is usually the whole of the difference that matters.

SAME LINE, TWICE
 
sent as-is     2026-03-02T22:55:12Z r.scott@ne.com NE-LEWIS-LT 4624 3 NTLM SUCCESS
sent redacted  TIMESTAMP USER1 HOST1 4624 3 NTLM SUCCESS
 
the answer     "Windows security event 4624, logon type 3 is a network logon,
               NTLM authentication, successful. This is a remote session
               against the host using NTLM rather than Kerberos."
 
identical from both. Every token the answer used is in the second line.

What is retained, for how long, and who can read it? Abuse monitoring is legitimate and it means a human at the vendor can potentially read a prompt. That is acceptable for category 1 and 2 content and is a genuine consideration for category 4.

Reference What to establish about a tool, once answer it before somebody needs it at 02:40
RETENTION      how long are prompts kept, and by whom
TRAINING       are prompts used to train, and can that be off
ACCESS         can a human at the vendor read a prompt, and when
JURISDICTION   where does the data sit, and does that matter here
TIER           does the answer change on the paid plan
 
Every one of these has a different answer per vendor, per tier
and per year, which is why it is a document rather than a memory.
 
The working assumption when it is unanswered: anything sent may
persist somewhere you cannot reach.
This is the artifact the M10 project asks for. Writing it once removes the question from every future shift.

Where does it run? Data residency is a real constraint for some organizations and irrelevant for others, and it is worth establishing which you are before the conversation rather than during it.

THE ANSWER YOU NEED, IN THE FORM AN ANALYST CAN USE
 
  SANCTIONED     [tool name], enterprise tier, no training on prompts
  FOR            log formats, query help, technique explanation
  NOT FOR        real hostnames, accounts, addresses, incident detail
  IF UNSURE      redact and send. If redaction breaks the question,
                 ask a person instead.
  WHERE          pinned in the SOC channel, and in the runbook
 
Five lines. Findable in twenty seconds at two in the morning.

Can you demonstrate the terms to an auditor? Not whether you believe the terms are adequate, but whether you can produce them. The question is asked after an incident and it is asked in writing.

Model quality is not on that list. Whether a tool is good and whether it is safe to send data to are separate decisions.

Conflating them is how organizations end up with a tool that is safe and useless, which analysts then route around.

07

If it has already happened

What to do, and what not to bother doing

It has. Not necessarily to you, and if you manage a team, to somebody on it. Treat that as the starting condition rather than the failure state.

The useful response is not disciplinary. It is to establish what category was disclosed, because that determines whether anything needs doing at all. Category 1 or 2 needs nothing. Category 3 is worth recording, because a naming convention disclosed once is disclosed permanently and it belongs in your threat model rather than in a disciplinary file. Category 4, during a live incident, is a decision for whoever is running the incident and probably for legal.

IT HAS ALREADY HAPPENED. WHAT ACTUALLY HELPS. DOES NOT HELP disciplinary response asking people to be careful trying to retract the prompt all three leave the cause in place HELPS establish which category it was rotate anything live sanction a tool, so the next person has somewhere to go The absence of a sanctioned tool is the finding. The paste is the symptom.

The three on the left leave the cause in place. Only the right-hand column changes what happens the next time somebody is stuck.

The failure mode to avoid is silence. An analyst who believes disclosure ends their career does not report it, which converts a small, containable event into an unknown one. Whatever the policy says, the operational goal is that people tell you, and that goal is set by how the first report is handled rather than by what is written down.

Keep this A disclosure has happened. What to do, in order
1  ESTABLISH THE CATEGORY   1 or 2, nothing to do
                            3, record it, it is now permanent
                            4, whoever runs the incident decides
 
2  ROTATE ANYTHING LIVE     credentials, tokens, keys in the paste
 
3  SANCTION A TOOL          the absence of one is the finding
 
DO NOT      run a disciplinary process
            ask people to be more careful
            try to retract the prompt
 
ALL THREE LEAVE THE CAUSE IN PLACE, AND THE THIRD DOES NOT WORK.
The operational goal is that people tell you, and that is set by how the first report is handled rather than by what the policy says.

Module 7 develops this into a working policy: what to sanction, how to write terms that survive a vendor change, and how to build a workflow where the compliant route is genuinely the quickest one available.

08

Why redaction does not cost you the answer

What the assistant actually reads, and why the substitution is nearly free

The objection to redacting is that it removes the thing the assistant needed, and understanding why it usually does not is what makes the habit stick.

A model reading an unfamiliar log line is doing structural work. This token is positional, that one is a timestamp, the third has the shape of an identifier, the delimiters suggest a particular vendor's format. Almost none of that depends on the values themselves. Replace every username with USER1 and the structural answer comes back identical, because the question was about shape.

FIFTEEN SECONDS, AND THE ANSWER IS UNCHANGED r.scott@ne.com NE-LEWIS-LT 10.0.1.142 your estate replace with USER1, HOST1, 10.0.0.1 4624 logon type 3 NTLM SUCCESS the answer The tokens that identify you and the tokens that answer the question are different tokens. That is why redaction is nearly free, and why it is worth building as a reflex.

The tokens that identify you and the tokens that answer the question are different tokens, which is why the substitution is nearly free.

THE ANSWER COMES FROM THE TOKENS THAT ARE SAFE TO SEND VENDOR VOCABULARY 4624 · NTLM · SUCCESS logon type 3 · error codes what the answer is built from SEND THESE YOUR ESTATE r.scott@ne.com NE-LEWIS-LT · 45.83.64.117 contributes nothing to the answer REPLACE THESE The split is almost perfect, which is why redaction costs you nothing.

Read the left column and ask which of those tokens you could not have supplied from a vendor manual.

DO IT ON THE LINE FROM THE OPENING PASTE. FIFTEEN SECONDS.
 
BEFORE  2026-03-02T22:55:12Z r.scott@ne.com NE-LEWIS-LT 45.83.64.117 4624 3 Ntlm SUCCESS
AFTER   TIMESTAMP           USER1          HOST1       10.0.0.1     4624 3 Ntlm SUCCESS
 
STILL THERE   the event ID, the logon type, the auth package,
              the outcome. Every token the answer is built from.
GONE          who you are, and how you name things.

Where values do matter, they are the ones that are safe to keep. A result code, an event ID, a protocol name, an error string: these are vendor vocabulary rather than estate detail, and they are precisely the tokens the assistant needs to recognize the format.

Replace consistently rather than deleting. The relationships between values are often the point: "the same user appears in lines 1, 4 and 9" survives redaction perfectly well if that user is USER1 throughout, and it is frequently the fact that mattered. Deleting the values destroys that; substituting them does not.

09

Practice

Redact one real log line, and time it
hands on

Reading this section changes nothing. Fifteen seconds of redaction on one real line is what turns it into a reflex, and the reflex is the only part that survives a busy shift.

Practice Redact one real log line, and time it
  1. Take a log line you would genuinely have pasted somewhere.
  2. Replace consistently: usernames to USER1, hostnames to HOST1, addresses to 10.0.0.1.
  3. Keep result codes, event IDs and error strings. Those are vendor vocabulary and the assistant needs them.
It takes about fifteen seconds and the answer you get back does not change.

Then find out what your organization has actually sanctioned, and whether the answer is written down anywhere an analyst could find it at two in the morning. An unanswered question resolves itself in favor of the shortcut, every time.

Next: section 0.7 is the drill. Four generated queries against live data, three of them wrong and one of them fine.