Lesson 3.5 · Static Triage· 50 min
Writing Your First YARA Rules
Turn triage findings into YARA rules: strings, hex patterns, the pe and math modules, performance, and testing against goodware and variants.
Objectives
- Explain where YARA rules run in a defensive pipeline and what a match does and does not prove
- Write rules with text, hex and regex strings, modifiers, and conditions using the pe and math modules
- Choose rule strings from triage findings that survive recompilation but not unrelated software
- Avoid slow rules by understanding atoms, short strings and unanchored regular expressions
- Test a rule for false positives on a goodware corpus and for true positives on variants, then version it
Everything you have done so far in this module produces observations: a hash, a handful of odd strings, an import table that hints at injection, a section with suspicious entropy. Observations about one file are useful for one incident. A detection rule turns them into something that finds the next file too, on every endpoint, in every sandbox run, across years of stored samples.
YARA is the standard language for that job. A YARA rule describes a pattern of bytes and structure, and the YARA engine answers a single question for each file or memory region it scans: does this match? This lesson teaches you to write rules that answer that question well: fast, specific to the thing you analysed, and robust to the small changes attackers make between builds.
Where YARA runs
A rule you write today can end up in many places:
| Where | What it scans | Why it matters for rule design |
|---|---|---|
| Endpoint and server scanners (EDR plug-ins, THOR, Loki, osquery's YARA table) | Files on disk, sometimes process memory | Runs across thousands of files per host, so speed matters |
| Sandboxes and mail gateways | Submitted files, dropped files, memory dumps | Used for automatic classification and tagging |
| VirusTotal Livehunt and Retrohunt | New uploads, or months of past uploads | Finds related samples you did not know existed |
Memory scanning (yara <rules> <pid>, Volatility's yarascan) | A live process or a memory image | Catches packed malware after it unpacks itself |
| Your own sample repository | Your collection | Clustering and "have we seen this before?" |
Two consequences follow. First, a rule must be cheap enough to run everywhere, which is why performance gets its own section below. Second, a YARA match is a pattern match, not a verdict. It says "these bytes are present", nothing about intent. A rule for a malware family is only as good as the uniqueness of the patterns you chose.
Anatomy of a rule
import "pe"
rule MAL_Win_Example_Loader
{
meta:
description = "Detects Example loader by its config marker and decoder"
author = "Analyst Name"
date = "2026-09-29"
version = "1.0"
strings:
$cfg = "CFGv2|" ascii
$ua = "ExampleAgent" ascii wide
$hex = { 8B 45 ?? 35 [2-4] 89 45 }
$re = /https?:\/\/[a-z0-9.-]{4,40}\/gate\.php/ ascii
condition:
uint16(0) == 0x5A4D and filesize < 1MB and 2 of them
}A rule has a name, and three sections:
metaholds key/value pairs for humans and tooling. The engine ignores them when matching, but every serious rule set relies on them.stringsdeclares the patterns, each named with$. A rule can have no strings at all if its condition only uses file structure.conditionis a Boolean expression. It is the only mandatory section, and the rule matches when it evaluates to true.
Rule names follow identifier rules (letters, digits, underscores, not starting with a digit) and must be unique in a compiled rule set.
Text strings and modifiers
A text string matches literally, case-sensitive and as ASCII by default. Modifiers change that:
| Modifier | Effect | Typical use |
|---|---|---|
ascii | Match the one-byte-per-character form (the default when alone) | C strings, config text |
wide | Match UTF-16LE: each ASCII character followed by 00 | Windows W APIs, .NET, L"..." literals |
ascii wide | Match either form | When you do not know how the next build stores it |
nocase | Case-insensitive | PDB paths, file names, HTTP headers |
fullword | Match only if delimited by non-alphanumeric characters | Short names that also occur inside longer words |
xor / base64 | Match single-byte XOR or base64 encodings of the string | Lightly encoded strings |
fullword deserves care: "evil" with fullword matches C:\evil\x.exe
but not devils. It also fails on Agent/1.0 if you searched for Agent/
with fullword, because the next character is a digit.
Warning:
wideonly means "UTF-16LE with ASCII characters". It does not handle genuine non-ASCII text. And neither modifier helps with strings the program builds at runtime, like stackstrings or XOR-encrypted strings with a multi-byte key. For those you target the decoding code instead, or scan memory after decoding.
Hex strings
Hex strings match raw bytes, and are how you match code or binary structures:
| Syntax | Meaning | Example |
|---|---|---|
?? | Any byte | { E8 ?? ?? ?? ?? } (a call rel32) |
?5, 4? | Nibble wildcard | { 4? 8B 05 } (any REX prefix) |
[n-m] | Jump: between n and m arbitrary bytes | { 6A 40 [0-8] FF 15 } |
( A | B ) | Alternatives | { 8D (40 | 50) 5A } |
~XX | Any byte except XX | { 85 C0 ~74 } |
Wildcard the parts the compiler or linker picks for you: relative call and jump offsets, RIP-relative displacements, stack offsets. Keep literal the parts the author chose: constants, XOR keys, magic values, the shape of a decoding loop.
Regular expressions
Regexes sit between /.../ and use a Perl-like syntax. They are the most
flexible string type and the easiest to make slow. A good regex starts with a
literal run the engine can index (/LabAgent\/[0-9]\.[0-9]/), not with a
character class or .*.
Conditions
Conditions combine string references with arithmetic and file inspection:
condition:
uint16(0) == 0x5A4D // "MZ" at offset 0
and uint32(uint32(0x3C)) == 0x4550 // "PE\0\0" at e_lfanew
and filesize < 2MB
and ( $cfg or 2 of ($s*) )
and #ua >= 2 // $ua occurs at least twice
and $hex in (0 .. 4096) // within the first 4 KBuint16(0) == 0x5A4D is the idiom you will write most often. uint16 reads a
little-endian 16-bit value, so the bytes 4D 5A ("MZ") read as 0x5A4D.
Putting the header check and filesize first lets the engine skip most
non-matching files cheaply and stops a PE rule from firing on a text file that
happens to quote your strings (your own triage report, for instance).
The set operators are what make rules robust: any of them, all of them,
2 of ($s*), any of ($a*, $b1). $s* expands to every string whose name
starts with $s, so naming strings by role ($s_ for strong strings, $w_
for weak ones, $code_ for code) keeps conditions readable.
The pe and math modules
Modules expose parsed file structure to conditions. The pe module turns the
headers you studied in PE Headers and
PE Sections into fields:
import "pe"
import "math"
rule SUSP_PE_Injection_Import_Triad
{
condition:
pe.imports("kernel32.dll", "VirtualAllocEx")
and pe.imports("kernel32.dll", "WriteProcessMemory")
and pe.imports("kernel32.dll", "CreateRemoteThread")
}
rule SUSP_PE_HighEntropy_Text
{
condition:
uint16(0) == 0x5A4D
and for any i in (0 .. pe.number_of_sections - 1) : (
pe.sections[i].name == ".text"
and math.entropy(pe.sections[i].raw_data_offset,
pe.sections[i].raw_data_size) >= 7.2
)
}
rule SUSP_PE_FewSections_NoText
{
condition:
uint16(0) == 0x5A4D
and pe.number_of_sections <= 3
and not for any s in pe.sections : ( s.name == ".text" )
}Useful fields and functions:
| Expression | Returns |
|---|---|
pe.imports("dll", "Func") | True if that function is imported (DLL name is case-insensitive) |
pe.imphash() | The import hash covered in Identifying and Hashing Files |
pe.number_of_sections | Section count from the file header |
pe.sections[i].name, .raw_data_offset, .raw_data_size | Per-section data |
pe.pdb_path | The PDB path from the CodeView debug record, if any |
pe.is_dll(), pe.timestamp, pe.entry_point | Header facts |
math.entropy(offset, size) | Shannon entropy of a byte range, 0.0 to 8.0 |
These are exactly the signals from Reading Capabilities from Imports and Detecting Packers and Entropy, expressed as rules. Rules like the three above are hunting rules: they flag suspicious traits, not a family, and will match some legitimate software. That is fine as long as their names and metadata say so.
Warning:
pe.imphash()is only as unique as the import table. In the lab below, the lab program and a one-lineprintfprogram built with the same mingw-w64 toolchain produce the same imphash, because both import only the C runtime. Imphash identifies a family when the author's own imports dominate the table, not for small programs.
Performance: atoms and why rules get slow
YARA does not search for every string in every position. At compile time it extracts from each string a short atom (up to four bytes) that must be present for the string to match, and feeds all atoms into one Aho-Corasick automaton. Scanning a file is a single pass over its bytes looking for atoms; the full string is only verified where an atom hits. Performance therefore depends on the quality of the atoms:
- Short or common strings are poor atoms. A 2-byte string, or a hex string
like
{ 8D ?? ?? }, gives an atom that hits constantly, and every hit costs a verification. YARA warns:string "$a" may slow down scanning. - Common bytes are poor atoms.
00 00 00 00,FF FF FF FF,90 90,CC CCoccur everywhere in PE files. - Regexes without a literal prefix are expensive.
/[a-z]+\/[0-9]\.[0-9]/draws the same warning;/LabAgent\/[0-9]\.[0-9]/does not. - Large jumps and many alternatives multiply work. Prefer
[0-16]over[0-512]and split wildly different variants into separate strings. - Conditions are evaluated after scanning, so a cheap condition does not
save the string scan. But
filesizeand header checks still reduce the expensive module parsing and loop evaluation.
Treat the compiler warning as a failing test. The lab shows it firing on a real rule, and the fix.
From triage findings to rule strings
The best strings are things the author chose and would find tedious to change:
| Triage finding | Good rule material? | Notes |
|---|---|---|
| Custom user-agent, mutex, pipe or event names | Yes | Hard-coded, often reused across versions |
| Config markers, format strings, command names | Yes | e.g. "cmd=%s&id=%08x", "LABCFG{" |
| PDB path from the debug directory | Yes, as supporting evidence | Leaks user names and project names; see PE Headers |
| Distinctive code: decoders, key schedules, constants | Yes, if wildcarded | Survives string changes; may break across compilers |
| Typos and odd phrases in messages | Yes | Surprisingly stable across builds |
| C2 domains and IPs | Weak | Rotate constantly; better as IOCs than rule strings |
| Library, compiler and runtime strings | No | Shared by every program built the same way |
| Windows API names | No, on their own | Use pe.imports() or a capa rule instead |
| Whole-file hashes | Not in YARA | Use hash IOCs; a rule should generalise |
Strings and Obfuscated Strings is where most of these candidates come from. For each one, ask two questions: would the next build still contain it? and would a legitimate program ever contain it? A string that passes both is a strong string: a few are enough to match on their own. A string that passes only the first is weak: useful only in combination.
That distinction becomes the condition. A common pattern is "three of the strong strings, or one strong string plus supporting evidence", which tolerates the author removing or changing one artefact without matching on generic evidence alone.
Naming, metadata and versioning
A rule set is a shared codebase. Conventions from widely used public rule sets (see Florian Roth's YARA style guide) work well:
- Prefix by purpose:
MAL_for a malware family,HKTL_for hack tools,SUSP_for suspicious traits,HUNT_for broad hunting rules not meant for blocking. - Then platform, family and distinguishing detail:
MAL_Win_LabImplant_Strings,MAL_Win_LabImplant_Decoder. - Minimum metadata:
description,author,date,reference(your report ID or a public write-up), one or morehashvalues of samples the rule was tested on, and a sharing label such astlp. - Version it: bump
versionandmodifiedon every change, keep rules in Git, and keep the test samples' hashes next to the rule, so anyone can re-run the true-positive tests.
Tip: Write the
descriptionas the answer to "what does a match mean?" A SOC analyst reading an alert at 3 a.m. should know whether it means "known malware family" or "unusual packer, worth a look".
Lab: detect the lab implant
You will build a harmless program that contains the kind of artefacts a
real implant carries (a user-agent, a mutex name, a pipe name in UTF-16, a
config marker, a PDB name and a tiny decoding loop) but only prints them. It
has no networking, no persistence and creates no mutex. You need
mingw-w64, UPX, and YARA (brew install yara, apt install yara, or the
Windows release from the VirusTotal/yara GitHub page; pip install yara-python gives you the Python bindings).
-
Create
lab_implant.c:c /* lab_implant.c - a harmless "implant" for YARA practice. * It does nothing but print strings that look like implant artefacts. */ #include <windows.h> #include <stdio.h> #ifndef LAB_VERSION #define LAB_VERSION "1.0" #endif #ifndef LAB_INTERVAL #define LAB_INTERVAL "3600" #endif static const char USER_AGENT[] = "LabAgent/" LAB_VERSION " (Windows NT 10.0; lab-build)"; static const char MUTEX_NAME[] = "Global\\LabDemoMutex"; static const wchar_t PIPE_NAME[] = L"\\\\.\\pipe\\lab-demo-pipe"; static const char CONFIG[] = "LABCFG{interval=" LAB_INTERVAL ";jitter=20;mode=demo}"; /* An "encoded" blob and a tiny decoder, so the binary has distinctive code. */ unsigned char encoded_blob[] = { 0x3e, 0x3e, 0x31, 0x32, 0x7f }; __attribute__((noinline)) void lab_decode(unsigned char *buf, size_t len) { for (size_t i = 0; i < len; i++) buf[i] ^= (unsigned char)(0x5A + i); } int main(void) { lab_decode(encoded_blob, sizeof encoded_blob); printf("user-agent : %s\n", USER_AGENT); printf("mutex name : %s\n", MUTEX_NAME); printf("pipe name : %ls\n", PIPE_NAME); printf("config : %s\n", CONFIG); printf("decoded : %.5s\n", (char *)encoded_blob); return 0; } -
Build two "versions", the way an author's releases differ: v1 unoptimised, v2 optimised with a new version number and interval.
--pdbmakes the linker write a CodeView debug record naming a PDB file, as MSVC does. MSVC stores the full build path (often with a user name in it); GNU ld stores the name you pass.bash mkdir -p build/v1 build/v2 x86_64-w64-mingw32-gcc -O0 -s -Wl,--pdb=build/v1/lab_implant.pdb \ -o build/v1/lab_implant.exe lab_implant.c x86_64-w64-mingw32-gcc -O2 -s -DLAB_VERSION='"1.1"' -DLAB_INTERVAL='"900"' \ -Wl,--pdb=build/v2/lab_implant.pdb -o build/v2/lab_implant.exe lab_implant.c -
Triage v1 as in the previous lessons:
strings -a, GNUstrings -elor FLOSS (for the UTF-16 pipe name), and your import and hashing scripts. Then locate the decoder in a disassembly. It is the loop that adds0x5A:bash x86_64-w64-mingw32-objdump -d build/v1/lab_implant.exe | grep -B8 -A8 '0x5a('In my v1 build the loop contains these bytes (objdump output, addresses trimmed):
text 44 8d 40 5a lea 0x5a(%rax),%r8d 48 8b 55 10 mov 0x10(%rbp),%rdx 48 8b 45 f8 mov -0x8(%rbp),%rax 48 01 d0 add %rdx,%rax 44 31 c1 xor %r8d,%ecx 89 ca mov %ecx,%edx 88 10 mov %dl,(%rax) -
Write the obvious first rule,
naive.yar, from the exact user-agent and the exact code bytes:text rule LabImplant_naive { strings: $ua = "LabAgent/1.0 (Windows NT 10.0; lab-build)" $code = { 44 8D 40 5A 48 8B 55 10 48 8B 45 F8 48 01 D0 44 31 C1 } condition: all of them }bash yara -s naive.yar build/v1/lab_implant.exe build/v2/lab_implant.exetext LabImplant_naive build/v1/lab_implant.exe 0x2200:$ua: LabAgent/1.0 (Windows NT 10.0; lab-build) 0x8bc:$code: 44 8D 40 5A 48 8B 55 10 48 8B 45 F8 48 01 D0 44 31 C1It matches v1 and misses v2 entirely. The version bump broke
$ua, and-O2rewrote the loop, so$codeis gone too. Your offsets may differ; the result will not. -
Generalise. Compare the v2 disassembly (the same
greponbuild/v2/lab_implant.exe): the optimiser produced8d 50 5a 30 14 01 48 83 c0 01(lea 0x5a(%rax),%edx; xor %dl,(%rcx,%rax,1); add $0x1,%rax). A first attempt to cover both loops with one pattern,{ 8D (40 | 50) 5A [0-12] (30 | 31) }, matches both builds, but compiling it with warnings enabled reportsstring "$decode" may slow down scanning: its best atom is three common bytes. Split it into two longer, per-compiler patterns instead. The result,lab_implant.yar:text import "pe" rule MAL_Win_LabImplant_Strings { meta: description = "Detects the benign LabImplant training program by its embedded artefacts" author = "Your Name" date = "2026-09-29" modified = "2026-09-29" version = "1.2" reference = "TR-2026-0929-01 (internal triage report)" hash1 = "<sha256 of your build/v1/lab_implant.exe>" hash2 = "<sha256 of your build/v2/lab_implant.exe>" tlp = "CLEAR" strings: $s_ua = /LabAgent\/[0-9]{1,2}\.[0-9]{1,2} \(Windows NT/ ascii $s_mutex = "Global\\LabDemoMutex" ascii wide $s_pipe = "\\\\.\\pipe\\lab-demo-pipe" wide $s_cfg = "LABCFG{interval=" ascii $pdb = "lab_implant.pdb" ascii nocase fullword $decode_o0 = { 44 8D 40 5A [8-16] 44 31 C1 89 CA 88 10 } // -O0 loop $decode_o2 = { 8D 50 5A 30 14 01 48 83 C0 01 } // -O2 loop condition: uint16(0) == 0x5A4D and filesize < 2MB and ( 3 of ($s_*) or ($pdb and 1 of ($s_*)) or (1 of ($decode_*) and 2 of ($s_*)) ) }Note the escaping: in a YARA text string,
\\is one backslash, so"\\\\.\\pipe\\..."matches the literal\\.\pipe\.... -
Test for true positives on both versions:
bash yara -s lab_implant.yar build/v1/lab_implant.exe build/v2/lab_implant.exetext MAL_Win_LabImplant_Strings build/v1/lab_implant.exe 0x2200:$s_ua: LabAgent/1.0 (Windows NT 0x2230:$s_mutex: Global\LabDemoMutex 0x2260:$s_pipe: \\x00\\x00.\x00\\x00p\x00i\x00p\x00e\x00\\x00l\x00a\x00b\x00-\x00d\x00e\x00m\x00o\x00-\x00p\x00i\x00p\x00e\x00 0x22a0:$s_cfg: LABCFG{interval= 0x2e34:$pdb: lab_implant.pdb 0x8bc:$decode_o0: 44 8D 40 5A 48 8B 55 10 48 8B 45 F8 48 01 D0 44 31 C1 89 CA 88 10 MAL_Win_LabImplant_Strings build/v2/lab_implant.exe 0x2500:$s_ua: LabAgent/1.1 (Windows NT 0x24d0:$s_mutex: Global\LabDemoMutex 0x24a0:$s_pipe: \\x00\\x00.\x00\\x00p\x00i\x00p\x00e\x00\\x00l\x00a\x00b\x00-\x00d\x00e\x00m\x00o\x00-\x00p\x00i\x00p\x00e\x00 0x2460:$s_cfg: LABCFG{interval= 0x3034:$pdb: lab_implant.pdb 0xa80:$decode_o2: 8D 50 5A 30 14 01 48 83 C0 01-sprints every matching string with its offset. Read it critically: the rule would still match if any one artefact disappeared. -
Test what packing does. Pack v1 and scan it with both your rule and the structural rules from the
pesection above (save them aspe_hunting.yar):bash upx -q -o build/v1/lab_implant_upx.exe build/v1/lab_implant.exe yara lab_implant.yar build/v1/lab_implant_upx.exe yara pe_hunting.yar build/v1/lab_implant_upx.exeThe string rule finds nothing: the strings are compressed. Only
SUSP_PE_FewSections_NoTextfires (UPX leaves sections namedUPX0,UPX1,UPX2). This is why file rules and memory scanning complement each other, and why Detecting Packers and Entropy comes before rule writing in triage. -
Test for false positives. On Windows, scan the system directory; on Linux or macOS, scan
/usr/binwith a copy of the rule whose header check is replaced bytrue(otherwise theMZtest trivially rejects every ELF or Mach-O file):bash # Windows (cmd, in the lab VM) yara -r lab_implant.yar C:\Windows\System32 2>nul # Linux / macOS sed 's/uint16(0) == 0x5A4D/true/' lab_implant.yar > lab_implant_nohdr.yar yara -r lab_implant_nohdr.yar /usr/bin 2>/dev/nullExpect no output. Any hit is a string you believed was unique that is not; open the file, find out why, and tighten the rule. On a real team, this step runs against a large goodware corpus automatically for every rule change.
-
Check the imphash claim from the warning above:
bash printf '#include <stdio.h>\nint main(int c, char **v){printf("argc=%%d %%s\\n", c, v[0]);return 0;}\n' > hello2.c x86_64-w64-mingw32-gcc -O2 -s -o hello2.exe hello2.c python3 -c "import pefile,sys; [print(f, pefile.PE(f).get_imphash()) for f in sys.argv[1:]]" \ build/v1/lab_implant.exe build/v2/lab_implant.exe hello2.exe
Questions to answer: Which single artefact would you remove from the
source to make the rule miss, and does the condition allow for that? Why did
$s_mutex match only in its ASCII form even though the rule asks for ascii wide? If the author moved all strings behind the decoder, which parts of the
rule would still work, and what would you add? Why is the imphash of
hello2.exe the same as the implant's, and what does that tell you about
imphash rules for small programs?
Key takeaways
- A YARA rule is
meta+strings+condition; a match means "the pattern is present", and the rule's name and description must say what that implies. - Use
ascii wide,nocase,fullword, wildcards, jumps and alternatives to absorb the differences you expect between builds, and nothing more. - Start PE rules with
uint16(0) == 0x5A4Dand afilesizebound; use thepeandmathmodules for structure, imports and entropy. - Build strings from author-chosen artefacts (names, markers, PDB paths,
decoder code) and combine them with
N of ($s*)so one change does not break the rule. - Heed the "may slow down scanning" warning: poor atoms and unanchored regexes cost every scan, everywhere.
- A rule is not done until it has matched every known variant, matched nothing in a goodware corpus, and been versioned with the hashes it was tested on.