YARA Rule Writing for DFIR

Write it. Measure it. Say what it proves.

A rule that matches its own sample proves nothing, a rule that returns nothing has four causes that report identically, and a zero false-positive score can come from a corpus that could never have punished the pattern. This course teaches YARA by running it: every figure came from a command against a stated population, and the measurement wins where it contradicts the documentation.

Included with Premium, from $19.99/month, or $179/year and save 25%. Preview the first module free, no account needed.
Your subscription also includes the Practice Hub: graded scenarios, forensic cases, query drills and the response playbooks.
View Pricing Take End of Course Exam → 10 CPE Credits

What you'll be able to do

✓Write a rule and know what each modifier costs, including the one that takes a rule to zero matches while passing the linter
✓Run a scan whose result means something: the denominator recorded, the bounds chosen, and the two-sided control that separates a clean estate from a broken rule
✓Choose patterns by what it would cost the author to change them, measured against a version family and a clean corpus together
✓Read a file with yr dump before writing any condition, and use the pe and elf modules where the strings are gone
✓Write rules for documents and encoded scripts, where a correct rule matching nothing is a property of the format rather than a mistake
✓Find the one rule costing most of a scan, tune with one change at a time, and diagnose a rule that has gone quiet
✓Deploy across both engines, choose a threshold for a surface, and share a rule without making a claim it cannot support
SEC409 | Premium tier | 7 modules across 3 phases | 8–10 hours at your own pace | 10 CPE credits

Course Syllabus

Every module and every lesson. The first three modules are open; the rest open on a click.

Download the full syllabus (PDF)

Course Orientation

Module 0Course OrientationCourse Preview

What SEC409 teaches: writing YARA rules that still fire next month, on files you have not seen, without burying the estate in false positives. Built on YARA-X, with the 4.x deployment reality the rules actually land in. Start here.

Show 7 lessonsHide lessons
  1. 0.10.1 Your First Scan, and What It Does Not Tell YouPreview
  2. 0.20.2 The Specimens You Will BuildPreview
  3. 0.30.3 The Corpus Your Rules Must Not MatchPreview
  4. 0.40.4 Where Your Rules Will RunPreview
  5. 0.50.5 What a Match Is WorthPreview
  6. 0.60.6 What the Seven Modules BuildPreview
  7. 0.7Module SummaryPreview

Phase 1: Writing and Running Rules

Module 1The Rule as a Claim

The YARA language: text strings and their modifiers, hex patterns with wildcards and jumps, regular expressions and their cost, the condition logic that turns candidate patterns into an assertion, and metadata that survives being read by somebody else.

Show 8 lessonsHide lessons
  1. 1.11.1 Anatomy of a Rule
  2. 1.21.2 Text Strings and Modifiers
  3. 1.31.3 Hex Patterns and Wildcards
  4. 1.41.4 Regular Expressions
  5. 1.51.5 The Condition Is the Rule
  6. 1.61.6 Metadata and the Writing Loop
  7. 1.7Module Summary
  8. 1.8Check My Knowledge
Module 2Running Scans

Operating yr scan properly: what a target is and why directory scanning is not recursive by default, reading output at scale, threads and timeouts and size bounds, the controls that prove a negative, rule sets and namespaces, and the hunt loop.

Show 8 lessonsHide lessons
  1. 2.12.1 Targets and Recursion
  2. 2.22.2 Reading What Comes Back
  3. 2.32.3 Scanning at Scale
  4. 2.42.4 Controls, and Proving a Negative
  5. 2.52.5 Rule Sets and Namespaces
  6. 2.62.6 The Hunt Loop
  7. 2.7Module Summary
  8. 2.8Check My Knowledge

Phase 2: Choosing What to Match

Module 3Choosing What to Match

Choosing patterns by what it would cost the author to change them: the durability question, both measurement columns, extraction as a workflow, the false friends your eye picks first, and the build-environment trap.

Show 8 lessonsHide lessons
  1. 3.13.1 The Durability Question
  2. 3.23.2 The Durability Ladder
  3. 3.33.3 Extracting Candidates
  4. 3.43.4 Patterns That Look Good and Are Not
  5. 3.53.5 The Build-Environment Trap
  6. 3.63.6 Choosing the Set
  7. 3.7Module Summary
  8. 3.8Check My Knowledge
Module 4Modules and Structure

Reading a file with yr dump before writing any condition: pe fields and the bitmask trap, imports and imphash, entropy baselines, signatures and the Rich header, and what structural rules actually cost.

Show 8 lessonsHide lessons
  1. 4.14.1 Reading a File With yr dump
  2. 4.24.2 Writing Conditions on pe Fields
  3. 4.34.3 Imports and Imphash
  4. 4.44.4 Sections and Entropy
  5. 4.54.5 Signatures and Resources
  6. 4.64.6 Structural Rules in Practice
  7. 4.7Module Summary
  8. 4.8Check My Knowledge
Module 5Non-Binary Formats

Formats where the binary model does not apply: why a correct rule returns nothing on a document or an encoded script, what each container keeps readable, webshell precision, ELF differences, and the three questions that transfer.

Show 6 lessonsHide lessons
  1. 5.15.2 Office Documents Are Archives
  2. 5.25.3 Scripts and Layered Encoding
  3. 5.35.5 ELF and the Linux Estate
  4. 5.45.6 Choosing the Approach
  5. 5.5Module Summary
  6. 5.6Check My Knowledge

Phase 3: Maintaining and Deploying

Module 6Testing, Tuning and Performance

What happens after a rule is written: finding which rule in a set dominates a scan, the tuning loop, baselines and regression, diagnosing a rule that has gone quiet, reviewing rules you did not write, and the maintenance cadence.

Show 8 lessonsHide lessons
  1. 6.16.1 Finding the Expensive Rule
  2. 6.26.2 The Tuning Loop
  3. 6.36.3 Regression and Baselines
  4. 6.46.4 When a Rule Stops Firing
  5. 6.56.5 Reviewing Rules You Did Not Write
  6. 6.66.6 The Maintenance Cadence
  7. 6.7Module Summary
  8. 6.8Check My Knowledge
Module 7Deploying and Hunting

What changes when a rule leaves your machine: the two engines and where they disagree, surfaces and thresholds, deploying a set, running a hunt, sharing rules honestly, and where the course leaves you.

Show 8 lessonsHide lessons
  1. 7.17.1 The Two Engines
  2. 7.27.2 Surfaces and Thresholds
  3. 7.37.3 Deploying a Set
  4. 7.47.4 Hunting With Rules
  5. 7.57.5 Sharing and Attribution
  6. 7.67.6 Where This Leaves You
  7. 7.7Module Summary
  8. 7.8Check My Knowledge

Phase 0: Course Resources

ResourcesCheatsheets

The lookup layer: five sheets covering patterns and modifiers, scanning and output, choosing patterns, structure and modules, and non-binary formats, each entry stating what to read and what it does not prove.

Show 5 lessonsHide lessons
  1. 1Patterns and Modifiers
  2. 2Scanning, Output and Controls
  3. 3Choosing and Measuring Patterns
  4. 4Structure and Modules
  5. 5Non-Binary Formats and Deployment
ResourcesCookbooks

Three runbooks: writing a rule from a sample end to end, running a hunt whose negative result means something, and reviewing a rule set you did not write.

Show 3 lessonsHide lessons
  1. 1Writing a Rule From a Sample
  2. 2Running a Hunt
  3. 3Reviewing a Set You Did Not Write
ResourcesLab Setup

Building the environment every measurement in this course depends on: the yr binary, two specimen families, a clean corpus, a PE corpus, the non-binary specimens, and a controls directory, with a verification script that fails loudly.

Show 1 lessonHide lessons
  1. 1Building the Lab
ResourcesWalkthroughs

Two cases worked end to end including the wrong turns: a packed PE where the obvious signal points the wrong way, and a document-and-script chain where a correct rule matches nothing.

Show 2 lessonsHide lessons
  1. 1A Packed PE
  2. 2A Document and a Script
ResourcesPlaybooks

Three procedures for the moments something fires: a rule hit and you have to decide what it establishes, a rule has gone quiet and you cannot tell why, and a set stopped producing hits after an update.

Show 3 lessonsHide lessons
  1. 1A Rule Hit
  2. 2A Rule Has Gone Quiet
  3. 3A Set Stopped Producing Hits
ResourcesOperational Reference

The whole course condensed to one operating page: the command for each stage, the figure it produced, and the one judgment that stage turns on.

ResourcesReferences and Further Reading

The tools, the documentation, the frameworks, and what to check first when something in the course behaves differently from the page.

Course Completion

CompletionCourse Exam

YARA Rule Writing end-of-course exam: a hunt that returns zero hits including on the host holding the sample, testing whether you can separate did not match from was never evaluated and diagnose three independent failures under incident pressure.

Show 1 lessonHide lessons
  1. 1Course Completion. YARA Rule Writing

Course overview

YARA is taught almost everywhere as a syntax tour, which produces practitioners who can write a rule and cannot say what it establishes. This course was built by running the tool: every figure on every page came from a command against a stated population, and where a measurement contradicted what the documentation implied, the measurement is what the course teaches. Learn how to:

✓ Write a rule and know what each modifier costs, including the one that takes a rule to zero matches while passing the linter
✓ Run a scan whose result means something, with the denominator recorded and the bounds chosen rather than inherited
✓ Choose patterns by what it would cost the author to change them, measured against a version family and a clean corpus together
✓ Read a file with yr dump before writing any condition, and reach the structure that survives when the strings are gone
✓ Recognize the four causes of an empty result, and the two-sided control that separates a clean estate from a broken rule

How this course works

A YARA rule is a claim about what a file contains. This course runs the same loop for every rule it writes, because a rule that matches your sample and nothing else is a hash with extra steps.

1. Decide what the rule is claiming. This family, this packer, this technique, or this specific sample. The four claims need different rules and have different lifespans.

2. Choose strings that survive a recompile. Anything the author would change between builds is a bad string. What the code has to do to work is a good one.

3. Bound it with structure. File size, format, section characteristics. A condition that reaches the strings only for plausible candidates is faster and matches less by accident.

4. Test against a corpus, not a sample. Run it over goodware as well as malware. A rule nobody ran against a clean corpus is a rule whose false positive rate is unknown rather than low.

5. Measure what it costs to run. Scanning speed is a deployment constraint. A rule too slow to run at scale does not protect anything.

What this course assumes

No minimum experience and no prerequisite course. File formats, hex, regular expressions and the PE structure are explained where they are first needed.

What makes it go faster: comfort with a hex editor and some exposure to file formats. Neither is required. Every rule in the course is built against samples the course provides.

What this course does not cover: reverse engineering, malware analysis as a discipline, and unpacking. Those decide what your rule should look for; this course is about writing the rule once you know.

Who this course is for

You are a detection engineer, threat hunter, incident responder, or the person handed a sample and asked whether the estate is affected. No prior YARA experience is assumed and every concept is explained where it is first used. This course is for you if you want to:

✓ Stop reading a zero false-positive score as precision without asking what the corpus could have punished
✓ Write a rule that survives the next build rather than one that matches the sample you extracted it from
✓ Say what a hit establishes, in a sentence that holds up in a report you are not present to defend
✓ Diagnose a rule that has gone quiet, rather than retiring it on silence alone
✓ Audit an imported rule set in about a minute per rule, with evidence instead of an impression

What you'll learn

Seven teachable modules across three phases, working from a single rule to a set you maintain and deploy.

✓ Rule anatomy and every modifier measured, with the condition read aloud as a claim about a file
✓ Scanning properly: targets and recursion, the four output formats, bounds and their cost in coverage
✓ Pattern selection as cost to the author, the two measurement columns, and the two ways a measurement misleads you
✓ Structure with the pe and elf modules, entropy against a baseline, imphash and where it collides
✓ Documents and encoded scripts, where a correct rule matching nothing is the format rather than a mistake
✓ Finding the rule costing most of a scan, tuning one change at a time, and diagnosing silence
✓ Deploying across both engines, choosing a threshold per surface, and sharing a rule honestly

Key course takeaways

✓ A rule is a claim about bytes, and everything a reader adds to a hit comes from somewhere else
✓ A measurement is a statement about the conditions it was taken under, which three separate modules arrive at independently
✓ Absence needs a two-sided check, because a negation proves the scan reached the files and never that the rule could fire
✓ A pattern built on a constraint outlives one built on a choice, which is the question to ask of every candidate
✓ Where a static file rule stops, named as four specific boundaries rather than left as a gap in your skill

Where this fits in your workflow

YARA sits between investigation and detection. In an incident you write a rule to find the same file across the rest of the estate. In a hunt you take somebody else's rule and have to decide what a hit from it establishes. In detection engineering you decide which surface a rule is precise enough for, which is a measurement rather than a preference.

It connects directly to Malware Triage, where a rule is one output of a verdict, Microsoft Cloud Incident Response and Windows Endpoint Investigation for scanning collected evidence, and Memory Forensics, which is where the packing boundary this course names finally dissolves.

What this course is not

It is not reverse engineering. You will not disassemble anything, and the course names the point where a static file rule has done what it can rather than pretending to reach past it. Reading strings output is the whole prerequisite.

It is not a syntax reference. The official documentation lists every keyword, and this course was built by running the tool against stated populations: which patterns survive a rebuild, what a false-positive figure is a statement about, and what a hit does and does not let you claim.

Things you need to know

What are the prerequisites for this course?

None. Every concept is explained where it is first used, and the only assumed skill is reading the output of strings on a binary. No reverse engineering, assembly or malware analysis experience is needed.

Do I need live malware to follow along?

No, and the course was built that way deliberately. Every specimen is buildable from a few lines of C and a compiler, which is what makes the measurements reproducible: you compile five builds of one program, change one thing per build, and watch which patterns die at which change. The clean corpus is the system binaries already on the machine.

Which YARA do I need?

YARA-X, the current Rust implementation, using its yr command line tool. Install YARA 4.x alongside it if you can, because one module measures eight constructs where the two engines disagree and compiling on both is what removes that class of problem.

Does this work on Windows or macOS?

The tool does. The lab as written assumes Linux, because the clean corpus is /usr/bin and the specimens are built with gcc. A Linux VM is enough, and the structural module works on Windows executables you supply.

Is this course current?

Every command and flag was executed against YARA-X 1.19.0 at the time of writing rather than taken from documentation, and in eight places the measurement contradicted what the documentation implied. Tool output formats change, so the references module names what to check first when something behaves differently from the page.

Usage rights and disclaimer

Course materials are licensed for your individual use. You may apply everything you build here in your own environment and in client work. You may not redistribute the course content itself or resell it as training.

Specimens, outputs and figures are illustrative and were measured against the populations each page states. Verify against your own tooling before relying on any specific behavior, and treat what your own commands return as authoritative over any document, including this one.

Rules written in this course are claims about bytes. What a hit establishes is bounded by the patterns it required, and no rule here attributes a file to an actor.

Ridgeline Cyber is not affiliated with any tool vendor named here. Product names are used descriptively.

COURSE ASSESSMENT

End of Course Exam

Complete the course, then prove your skills under time pressure. Pass mark: 70. Earn your certificate with CPE credits.

40minutes
3phases
100points
1scenario
Take End of Course Exam

One random scenario per attempt. Certificate issued on pass.