Skip to content

Leçon 3.5 · Triage statique· 50 min

Writing Your First YARA Rules

Turn triage findings into YARA rules: strings, hex patterns, the pe and math modules, performance, and testing against goodware and variants.

Cette leçon n’est disponible qu’en anglais pour le moment.

Objectifs

  • Explain where YARA rules run in a defensive pipeline and what a match does and does not prove
  • Write rules with text, hex and regex strings, modifiers, and conditions using the pe and math modules
  • Choose rule strings from triage findings that survive recompilation but not unrelated software
  • Avoid slow rules by understanding atoms, short strings and unanchored regular expressions
  • Test a rule for false positives on a goodware corpus and for true positives on variants, then version it

Everything you have done so far in this module produces observations: a hash, a handful of odd strings, an import table that hints at injection, a section with suspicious entropy. Observations about one file are useful for one incident. A detection rule turns them into something that finds the next file too, on every endpoint, in every sandbox run, across years of stored samples.

YARA is the standard language for that job. A YARA rule describes a pattern of bytes and structure, and the YARA engine answers a single question for each file or memory region it scans: does this match? This lesson teaches you to write rules that answer that question well: fast, specific to the thing you analysed, and robust to the small changes attackers make between builds.

Where YARA runs

A rule you write today can end up in many places:

WhereWhat it scansWhy it matters for rule design
Endpoint and server scanners (EDR plug-ins, THOR, Loki, osquery's YARA table)Files on disk, sometimes process memoryRuns across thousands of files per host, so speed matters
Sandboxes and mail gatewaysSubmitted files, dropped files, memory dumpsUsed for automatic classification and tagging
VirusTotal Livehunt and RetrohuntNew uploads, or months of past uploadsFinds related samples you did not know existed
Memory scanning (yara <rules> <pid>, Volatility's yarascan)A live process or a memory imageCatches packed malware after it unpacks itself
Your own sample repositoryYour collectionClustering and "have we seen this before?"

Two consequences follow. First, a rule must be cheap enough to run everywhere, which is why performance gets its own section below. Second, a YARA match is a pattern match, not a verdict. It says "these bytes are present", nothing about intent. A rule for a malware family is only as good as the uniqueness of the patterns you chose.

Anatomy of a rule

text
import "pe"

rule MAL_Win_Example_Loader
{
    meta:
        description = "Detects Example loader by its config marker and decoder"
        author      = "Analyst Name"
        date        = "2026-09-29"
        version     = "1.0"

    strings:
        $cfg  = "CFGv2|" ascii
        $ua   = "ExampleAgent" ascii wide
        $hex  = { 8B 45 ?? 35 [2-4] 89 45 }
        $re   = /https?:\/\/[a-z0-9.-]{4,40}\/gate\.php/ ascii

    condition:
        uint16(0) == 0x5A4D and filesize < 1MB and 2 of them
}

A rule has a name, and three sections:

  • meta holds key/value pairs for humans and tooling. The engine ignores them when matching, but every serious rule set relies on them.
  • strings declares the patterns, each named with $. A rule can have no strings at all if its condition only uses file structure.
  • condition is a Boolean expression. It is the only mandatory section, and the rule matches when it evaluates to true.

Rule names follow identifier rules (letters, digits, underscores, not starting with a digit) and must be unique in a compiled rule set.

Text strings and modifiers

A text string matches literally, case-sensitive and as ASCII by default. Modifiers change that:

ModifierEffectTypical use
asciiMatch the one-byte-per-character form (the default when alone)C strings, config text
wideMatch UTF-16LE: each ASCII character followed by 00Windows W APIs, .NET, L"..." literals
ascii wideMatch either formWhen you do not know how the next build stores it
nocaseCase-insensitivePDB paths, file names, HTTP headers
fullwordMatch only if delimited by non-alphanumeric charactersShort names that also occur inside longer words
xor / base64Match single-byte XOR or base64 encodings of the stringLightly encoded strings

fullword deserves care: "evil" with fullword matches C:\evil\x.exe but not devils. It also fails on Agent/1.0 if you searched for Agent/ with fullword, because the next character is a digit.

Warning: wide only means "UTF-16LE with ASCII characters". It does not handle genuine non-ASCII text. And neither modifier helps with strings the program builds at runtime, like stackstrings or XOR-encrypted strings with a multi-byte key. For those you target the decoding code instead, or scan memory after decoding.

Hex strings

Hex strings match raw bytes, and are how you match code or binary structures:

SyntaxMeaningExample
??Any byte{ E8 ?? ?? ?? ?? } (a call rel32)
?5, 4?Nibble wildcard{ 4? 8B 05 } (any REX prefix)
[n-m]Jump: between n and m arbitrary bytes{ 6A 40 [0-8] FF 15 }
( A | B )Alternatives{ 8D (40 | 50) 5A }
~XXAny byte except XX{ 85 C0 ~74 }

Wildcard the parts the compiler or linker picks for you: relative call and jump offsets, RIP-relative displacements, stack offsets. Keep literal the parts the author chose: constants, XOR keys, magic values, the shape of a decoding loop.

Regular expressions

Regexes sit between /.../ and use a Perl-like syntax. They are the most flexible string type and the easiest to make slow. A good regex starts with a literal run the engine can index (/LabAgent\/[0-9]\.[0-9]/), not with a character class or .*.

Conditions

Conditions combine string references with arithmetic and file inspection:

text
condition:
    uint16(0) == 0x5A4D                // "MZ" at offset 0
    and uint32(uint32(0x3C)) == 0x4550 // "PE\0\0" at e_lfanew
    and filesize < 2MB
    and ( $cfg or 2 of ($s*) )
    and #ua >= 2                       // $ua occurs at least twice
    and $hex in (0 .. 4096)            // within the first 4 KB

uint16(0) == 0x5A4D is the idiom you will write most often. uint16 reads a little-endian 16-bit value, so the bytes 4D 5A ("MZ") read as 0x5A4D. Putting the header check and filesize first lets the engine skip most non-matching files cheaply and stops a PE rule from firing on a text file that happens to quote your strings (your own triage report, for instance).

The set operators are what make rules robust: any of them, all of them, 2 of ($s*), any of ($a*, $b1). $s* expands to every string whose name starts with $s, so naming strings by role ($s_ for strong strings, $w_ for weak ones, $code_ for code) keeps conditions readable.

The pe and math modules

Modules expose parsed file structure to conditions. The pe module turns the headers you studied in PE Headers and PE Sections into fields:

text
import "pe"
import "math"

rule SUSP_PE_Injection_Import_Triad
{
    condition:
        pe.imports("kernel32.dll", "VirtualAllocEx")
        and pe.imports("kernel32.dll", "WriteProcessMemory")
        and pe.imports("kernel32.dll", "CreateRemoteThread")
}

rule SUSP_PE_HighEntropy_Text
{
    condition:
        uint16(0) == 0x5A4D
        and for any i in (0 .. pe.number_of_sections - 1) : (
            pe.sections[i].name == ".text"
            and math.entropy(pe.sections[i].raw_data_offset,
                             pe.sections[i].raw_data_size) >= 7.2
        )
}

rule SUSP_PE_FewSections_NoText
{
    condition:
        uint16(0) == 0x5A4D
        and pe.number_of_sections <= 3
        and not for any s in pe.sections : ( s.name == ".text" )
}

Useful fields and functions:

ExpressionReturns
pe.imports("dll", "Func")True if that function is imported (DLL name is case-insensitive)
pe.imphash()The import hash covered in Identifying and Hashing Files
pe.number_of_sectionsSection count from the file header
pe.sections[i].name, .raw_data_offset, .raw_data_sizePer-section data
pe.pdb_pathThe PDB path from the CodeView debug record, if any
pe.is_dll(), pe.timestamp, pe.entry_pointHeader facts
math.entropy(offset, size)Shannon entropy of a byte range, 0.0 to 8.0

These are exactly the signals from Reading Capabilities from Imports and Detecting Packers and Entropy, expressed as rules. Rules like the three above are hunting rules: they flag suspicious traits, not a family, and will match some legitimate software. That is fine as long as their names and metadata say so.

Warning: pe.imphash() is only as unique as the import table. In the lab below, the lab program and a one-line printf program built with the same mingw-w64 toolchain produce the same imphash, because both import only the C runtime. Imphash identifies a family when the author's own imports dominate the table, not for small programs.

Performance: atoms and why rules get slow

YARA does not search for every string in every position. At compile time it extracts from each string a short atom (up to four bytes) that must be present for the string to match, and feeds all atoms into one Aho-Corasick automaton. Scanning a file is a single pass over its bytes looking for atoms; the full string is only verified where an atom hits. Performance therefore depends on the quality of the atoms:

  • Short or common strings are poor atoms. A 2-byte string, or a hex string like { 8D ?? ?? }, gives an atom that hits constantly, and every hit costs a verification. YARA warns: string "$a" may slow down scanning.
  • Common bytes are poor atoms. 00 00 00 00, FF FF FF FF, 90 90, CC CC occur everywhere in PE files.
  • Regexes without a literal prefix are expensive. /[a-z]+\/[0-9]\.[0-9]/ draws the same warning; /LabAgent\/[0-9]\.[0-9]/ does not.
  • Large jumps and many alternatives multiply work. Prefer [0-16] over [0-512] and split wildly different variants into separate strings.
  • Conditions are evaluated after scanning, so a cheap condition does not save the string scan. But filesize and header checks still reduce the expensive module parsing and loop evaluation.

Treat the compiler warning as a failing test. The lab shows it firing on a real rule, and the fix.

From triage findings to rule strings

The best strings are things the author chose and would find tedious to change:

Triage findingGood rule material?Notes
Custom user-agent, mutex, pipe or event namesYesHard-coded, often reused across versions
Config markers, format strings, command namesYese.g. "cmd=%s&id=%08x", "LABCFG{"
PDB path from the debug directoryYes, as supporting evidenceLeaks user names and project names; see PE Headers
Distinctive code: decoders, key schedules, constantsYes, if wildcardedSurvives string changes; may break across compilers
Typos and odd phrases in messagesYesSurprisingly stable across builds
C2 domains and IPsWeakRotate constantly; better as IOCs than rule strings
Library, compiler and runtime stringsNoShared by every program built the same way
Windows API namesNo, on their ownUse pe.imports() or a capa rule instead
Whole-file hashesNot in YARAUse hash IOCs; a rule should generalise

Strings and Obfuscated Strings is where most of these candidates come from. For each one, ask two questions: would the next build still contain it? and would a legitimate program ever contain it? A string that passes both is a strong string: a few are enough to match on their own. A string that passes only the first is weak: useful only in combination.

That distinction becomes the condition. A common pattern is "three of the strong strings, or one strong string plus supporting evidence", which tolerates the author removing or changing one artefact without matching on generic evidence alone.

Naming, metadata and versioning

A rule set is a shared codebase. Conventions from widely used public rule sets (see Florian Roth's YARA style guide) work well:

  • Prefix by purpose: MAL_ for a malware family, HKTL_ for hack tools, SUSP_ for suspicious traits, HUNT_ for broad hunting rules not meant for blocking.
  • Then platform, family and distinguishing detail: MAL_Win_LabImplant_Strings, MAL_Win_LabImplant_Decoder.
  • Minimum metadata: description, author, date, reference (your report ID or a public write-up), one or more hash values of samples the rule was tested on, and a sharing label such as tlp.
  • Version it: bump version and modified on every change, keep rules in Git, and keep the test samples' hashes next to the rule, so anyone can re-run the true-positive tests.

Tip: Write the description as the answer to "what does a match mean?" A SOC analyst reading an alert at 3 a.m. should know whether it means "known malware family" or "unusual packer, worth a look".

Lab: detect the lab implant

You will build a harmless program that contains the kind of artefacts a real implant carries (a user-agent, a mutex name, a pipe name in UTF-16, a config marker, a PDB name and a tiny decoding loop) but only prints them. It has no networking, no persistence and creates no mutex. You need mingw-w64, UPX, and YARA (brew install yara, apt install yara, or the Windows release from the VirusTotal/yara GitHub page; pip install yara-python gives you the Python bindings).

  1. Create lab_implant.c:

    c
    /* lab_implant.c - a harmless "implant" for YARA practice.
     * It does nothing but print strings that look like implant artefacts. */
    #include <windows.h>
    #include <stdio.h>
    
    #ifndef LAB_VERSION
    #define LAB_VERSION "1.0"
    #endif
    #ifndef LAB_INTERVAL
    #define LAB_INTERVAL "3600"
    #endif
    
    static const char    USER_AGENT[] = "LabAgent/" LAB_VERSION " (Windows NT 10.0; lab-build)";
    static const char    MUTEX_NAME[] = "Global\\LabDemoMutex";
    static const wchar_t PIPE_NAME[]  = L"\\\\.\\pipe\\lab-demo-pipe";
    static const char    CONFIG[]     = "LABCFG{interval=" LAB_INTERVAL ";jitter=20;mode=demo}";
    
    /* An "encoded" blob and a tiny decoder, so the binary has distinctive code. */
    unsigned char encoded_blob[] = { 0x3e, 0x3e, 0x31, 0x32, 0x7f };
    
    __attribute__((noinline))
    void lab_decode(unsigned char *buf, size_t len) {
        for (size_t i = 0; i < len; i++)
            buf[i] ^= (unsigned char)(0x5A + i);
    }
    
    int main(void) {
        lab_decode(encoded_blob, sizeof encoded_blob);
        printf("user-agent : %s\n", USER_AGENT);
        printf("mutex name : %s\n", MUTEX_NAME);
        printf("pipe name  : %ls\n", PIPE_NAME);
        printf("config     : %s\n", CONFIG);
        printf("decoded    : %.5s\n", (char *)encoded_blob);
        return 0;
    }
  2. Build two "versions", the way an author's releases differ: v1 unoptimised, v2 optimised with a new version number and interval. --pdb makes the linker write a CodeView debug record naming a PDB file, as MSVC does. MSVC stores the full build path (often with a user name in it); GNU ld stores the name you pass.

    bash
    mkdir -p build/v1 build/v2
    x86_64-w64-mingw32-gcc -O0 -s -Wl,--pdb=build/v1/lab_implant.pdb \
        -o build/v1/lab_implant.exe lab_implant.c
    x86_64-w64-mingw32-gcc -O2 -s -DLAB_VERSION='"1.1"' -DLAB_INTERVAL='"900"' \
        -Wl,--pdb=build/v2/lab_implant.pdb -o build/v2/lab_implant.exe lab_implant.c
  3. Triage v1 as in the previous lessons: strings -a, GNU strings -el or FLOSS (for the UTF-16 pipe name), and your import and hashing scripts. Then locate the decoder in a disassembly. It is the loop that adds 0x5A:

    bash
    x86_64-w64-mingw32-objdump -d build/v1/lab_implant.exe | grep -B8 -A8 '0x5a('

    In my v1 build the loop contains these bytes (objdump output, addresses trimmed):

    text
    44 8d 40 5a     lea    0x5a(%rax),%r8d
    48 8b 55 10     mov    0x10(%rbp),%rdx
    48 8b 45 f8     mov    -0x8(%rbp),%rax
    48 01 d0        add    %rdx,%rax
    44 31 c1        xor    %r8d,%ecx
    89 ca           mov    %ecx,%edx
    88 10           mov    %dl,(%rax)
  4. Write the obvious first rule, naive.yar, from the exact user-agent and the exact code bytes:

    text
    rule LabImplant_naive
    {
        strings:
            $ua   = "LabAgent/1.0 (Windows NT 10.0; lab-build)"
            $code = { 44 8D 40 5A 48 8B 55 10 48 8B 45 F8 48 01 D0 44 31 C1 }
    
        condition:
            all of them
    }
    bash
    yara -s naive.yar build/v1/lab_implant.exe build/v2/lab_implant.exe
    text
    LabImplant_naive build/v1/lab_implant.exe
    0x2200:$ua: LabAgent/1.0 (Windows NT 10.0; lab-build)
    0x8bc:$code: 44 8D 40 5A 48 8B 55 10 48 8B 45 F8 48 01 D0 44 31 C1

    It matches v1 and misses v2 entirely. The version bump broke $ua, and -O2 rewrote the loop, so $code is gone too. Your offsets may differ; the result will not.

  5. Generalise. Compare the v2 disassembly (the same grep on build/v2/lab_implant.exe): the optimiser produced 8d 50 5a 30 14 01 48 83 c0 01 (lea 0x5a(%rax),%edx; xor %dl,(%rcx,%rax,1); add $0x1,%rax). A first attempt to cover both loops with one pattern, { 8D (40 | 50) 5A [0-12] (30 | 31) }, matches both builds, but compiling it with warnings enabled reports string "$decode" may slow down scanning: its best atom is three common bytes. Split it into two longer, per-compiler patterns instead. The result, lab_implant.yar:

    text
    import "pe"
    
    rule MAL_Win_LabImplant_Strings
    {
        meta:
            description = "Detects the benign LabImplant training program by its embedded artefacts"
            author      = "Your Name"
            date        = "2026-09-29"
            modified    = "2026-09-29"
            version     = "1.2"
            reference   = "TR-2026-0929-01 (internal triage report)"
            hash1       = "<sha256 of your build/v1/lab_implant.exe>"
            hash2       = "<sha256 of your build/v2/lab_implant.exe>"
            tlp         = "CLEAR"
    
        strings:
            $s_ua      = /LabAgent\/[0-9]{1,2}\.[0-9]{1,2} \(Windows NT/ ascii
            $s_mutex   = "Global\\LabDemoMutex" ascii wide
            $s_pipe    = "\\\\.\\pipe\\lab-demo-pipe" wide
            $s_cfg     = "LABCFG{interval=" ascii
            $pdb       = "lab_implant.pdb" ascii nocase fullword
            $decode_o0 = { 44 8D 40 5A [8-16] 44 31 C1 89 CA 88 10 }  // -O0 loop
            $decode_o2 = { 8D 50 5A 30 14 01 48 83 C0 01 }             // -O2 loop
    
        condition:
            uint16(0) == 0x5A4D
            and filesize < 2MB
            and (
                3 of ($s_*)
                or ($pdb and 1 of ($s_*))
                or (1 of ($decode_*) and 2 of ($s_*))
            )
    }

    Note the escaping: in a YARA text string, \\ is one backslash, so "\\\\.\\pipe\\..." matches the literal \\.\pipe\....

  6. Test for true positives on both versions:

    bash
    yara -s lab_implant.yar build/v1/lab_implant.exe build/v2/lab_implant.exe
    text
    MAL_Win_LabImplant_Strings build/v1/lab_implant.exe
    0x2200:$s_ua: LabAgent/1.0 (Windows NT
    0x2230:$s_mutex: Global\LabDemoMutex
    0x2260:$s_pipe: \\x00\\x00.\x00\\x00p\x00i\x00p\x00e\x00\\x00l\x00a\x00b\x00-\x00d\x00e\x00m\x00o\x00-\x00p\x00i\x00p\x00e\x00
    0x22a0:$s_cfg: LABCFG{interval=
    0x2e34:$pdb: lab_implant.pdb
    0x8bc:$decode_o0: 44 8D 40 5A 48 8B 55 10 48 8B 45 F8 48 01 D0 44 31 C1 89 CA 88 10
    MAL_Win_LabImplant_Strings build/v2/lab_implant.exe
    0x2500:$s_ua: LabAgent/1.1 (Windows NT
    0x24d0:$s_mutex: Global\LabDemoMutex
    0x24a0:$s_pipe: \\x00\\x00.\x00\\x00p\x00i\x00p\x00e\x00\\x00l\x00a\x00b\x00-\x00d\x00e\x00m\x00o\x00-\x00p\x00i\x00p\x00e\x00
    0x2460:$s_cfg: LABCFG{interval=
    0x3034:$pdb: lab_implant.pdb
    0xa80:$decode_o2: 8D 50 5A 30 14 01 48 83 C0 01

    -s prints every matching string with its offset. Read it critically: the rule would still match if any one artefact disappeared.

  7. Test what packing does. Pack v1 and scan it with both your rule and the structural rules from the pe section above (save them as pe_hunting.yar):

    bash
    upx -q -o build/v1/lab_implant_upx.exe build/v1/lab_implant.exe
    yara lab_implant.yar build/v1/lab_implant_upx.exe
    yara pe_hunting.yar build/v1/lab_implant_upx.exe

    The string rule finds nothing: the strings are compressed. Only SUSP_PE_FewSections_NoText fires (UPX leaves sections named UPX0, UPX1, UPX2). This is why file rules and memory scanning complement each other, and why Detecting Packers and Entropy comes before rule writing in triage.

  8. Test for false positives. On Windows, scan the system directory; on Linux or macOS, scan /usr/bin with a copy of the rule whose header check is replaced by true (otherwise the MZ test trivially rejects every ELF or Mach-O file):

    bash
    # Windows (cmd, in the lab VM)
    yara -r lab_implant.yar C:\Windows\System32 2>nul
    # Linux / macOS
    sed 's/uint16(0) == 0x5A4D/true/' lab_implant.yar > lab_implant_nohdr.yar
    yara -r lab_implant_nohdr.yar /usr/bin 2>/dev/null

    Expect no output. Any hit is a string you believed was unique that is not; open the file, find out why, and tighten the rule. On a real team, this step runs against a large goodware corpus automatically for every rule change.

  9. Check the imphash claim from the warning above:

    bash
    printf '#include <stdio.h>\nint main(int c, char **v){printf("argc=%%d %%s\\n", c, v[0]);return 0;}\n' > hello2.c
    x86_64-w64-mingw32-gcc -O2 -s -o hello2.exe hello2.c
    python3 -c "import pefile,sys; [print(f, pefile.PE(f).get_imphash()) for f in sys.argv[1:]]" \
        build/v1/lab_implant.exe build/v2/lab_implant.exe hello2.exe

Questions to answer: Which single artefact would you remove from the source to make the rule miss, and does the condition allow for that? Why did $s_mutex match only in its ASCII form even though the rule asks for ascii wide? If the author moved all strings behind the decoder, which parts of the rule would still work, and what would you add? Why is the imphash of hello2.exe the same as the implant's, and what does that tell you about imphash rules for small programs?

Key takeaways

  • A YARA rule is meta + strings + condition; a match means "the pattern is present", and the rule's name and description must say what that implies.
  • Use ascii wide, nocase, fullword, wildcards, jumps and alternatives to absorb the differences you expect between builds, and nothing more.
  • Start PE rules with uint16(0) == 0x5A4D and a filesize bound; use the pe and math modules for structure, imports and entropy.
  • Build strings from author-chosen artefacts (names, markers, PDB paths, decoder code) and combine them with N of ($s*) so one change does not break the rule.
  • Heed the "may slow down scanning" warning: poor atoms and unanchored regexes cost every scan, everywhere.
  • A rule is not done until it has matched every known variant, matched nothing in a goodware corpus, and been versioned with the hashes it was tested on.