Skip to content

Lesson 8.4 · Encoding, Crypto & Signatures· 55 min

Extracting Malware Configurations

Turn a sample's embedded configuration into normalised JSON with a locate-decode-parse-output extractor, and share the results responsibly.

Objectives

  • Explain what a malware configuration holds and why it is the highest-value output of an analysis
  • Enumerate where configurations live: resources, overlays, .data blobs, .NET settings, the registry and runtime fetches
  • Build an extractor on the locate–decode–parse–output pattern that emits normalised JSON
  • Describe how MWCP, CAPE extractors and RATDecoders-style projects organise extractors, and package one for MWCP
  • Make an extractor robust to version drift and share configurations responsibly

Of all the outputs an analysis produces, the configuration is the one the rest of the team wants first. A working triage report says what a sample is; the config says who it talks to and how. Extract it in an hour and you hand the SOC a set of live indicators and the detection engineers something to hunt on, long before a full reversing effort finishes. This lesson is about producing that config as clean, structured data — and about writing an extractor that keeps producing it as the family evolves.

It builds on the two skills either side of it in Module 8: recognising encodings and identifying crypto tell you how the config is protected, and scripting string decryption gives you the argument-recovery and reimplementation techniques you reuse here on a larger structure.

What a configuration holds

A malware family's config is the set of values that change from campaign to campaign while the code stays the same. Pulling them out tells you almost everything operationally important:

FieldWhy it matters
C2 hosts and portsNetwork IOCs; the core of any blocklist or hunt
Fallback / backup C2 listsInfrastructure you would otherwise miss
Campaign or botnet IDClusters samples into campaigns; attribution pivot
Encryption keys / RC4 keysDecrypt the C2 traffic and later payloads
Sleep / beacon intervalBehavioural signature; tunes detection windows
Mutex namesHost IOC; a reliable, cheap detection
Install paths and filenamesRemediation; where to clean
Registry keysPersistence locations to remove
Feature flagsWhat the sample is configured to actually do
Version stringTracks the family's development over time

Read that list against the output table from What Is Malware Analysis?: a single extracted config feeds IOCs, detection content, remediation guidance and intelligence at once. That is why config extraction, not full reversing, is where most analytical value per hour is created.

Where configurations live

Before you can decode a config you have to find it. Families hide it in a handful of predictable places:

  • PE resources. The .rsrc directory is a favourite: a config sits in a custom-named resource, often encrypted, sometimes disguised as an icon or bitmap. The resources and overlays lesson covers how to enumerate them.
  • Overlays. Data appended after the last section, past the end of the PE as the headers describe it, is invisible to naive parsers and a common config hiding place.
  • .data / .rdata blobs. A byte array in a data section, decrypted at startup. Usually preceded or wrapped by a marker — a magic string or byte sequence the code searches for. Markers are the single most useful thing an extractor can key on.
  • .NET settings classes. As covered in Analysing .NET Malware, managed families concentrate the config in a static settings class assigned in the .cctor, often Base64 then AES.
  • The registry or a dropped file. Some families write their config to the host at install and read it back, so the config is not in the original sample at all — you recover it from the infected host or a memory dump.
  • Fetched at runtime. The worst case: the sample downloads its config from C2 on first run. Static extraction gets you the decryption key and the URL; the config itself needs a live or simulated C2.

High entropy in a resource, an overlay or a .data blob is the tell that one of these holds an encrypted config.

The extractor pattern

Every config extractor, however different the family, has the same four stages. Keeping them separate is what makes an extractor readable, testable and maintainable.

  1. Locate. Find the config bytes. A YARA rule on the marker or on the decoding code is the standard mechanism, because it survives minor rebuilds and gives you the offset. Failing a marker, locate the resource by name, the overlay by parsing the PE, or the blob by an xref from the decode routine.
  2. Decode. Reverse the protection. This is where the encoding and crypto skills from earlier in the module pay off: reimplement the XOR, RC4, AES or custom scheme exactly, reusing the key you recovered from the code.
  3. Parse. Impose structure on the decoded bytes. Configs are typed: a length byte then a host, a fixed-offset table of fields, a length-prefixed list, a serialized blob (Protocol Buffers, JSON, a .NET BinaryFormatter stream). Read each field with its real type.
  4. Output. Emit normalised data — JSON with stable field names — so every sample of the family produces the same shape regardless of which fields are present. Normalisation is what lets downstream tooling consume configs from many families uniformly.

Tip: Write the four stages as four functions even for a one-off. When the next variant changes only the key, you edit decode and nothing else; when it adds a field, you edit parse. An extractor that mixes locating and parsing in one loop has to be rewritten for every variant.

Existing frameworks

You rarely start from a blank file. Several open frameworks give you the plumbing — sample loading, crypto helpers, a config schema, a runner — so your code is just the family-specific locate, decode and parse.

ProjectMaintainerShapeWhere it fits
DC3-MWCPDoD Cyber Crime CenterFramework where each parser is a Python class emitting a standard report of typed fieldsStandardised extractor output across many families
CAPE extractorsOpen-source (kevoreilly)Per-family Python modules run automatically against dumped payloads in the CAPE sandboxExtraction integrated with unpacking and detonation
malduckCERT PolskaLibrary of crypto, PE and structure helpers, plus a YARA-driven extractor modelBuilding blocks and a modern extractor runner
RATDecoders-style projectsCommunity (e.g. StormShield, Malwoverview lineage)A directory of standalone per-family scripts, each self-containedReading how a specific family is parsed; quick reference

They are organised the same way underneath: a registry of per-family extractors, each declaring how to recognise its family (usually a YARA rule) and how to produce a config, plus a runner that matches a sample to an extractor and normalises the result. MWCP calls the unit a parser; CAPE calls it an extractor module; malduck calls it an extractor class with an embedded yara_rule. Learning one transfers to the others.

Packaging your extractor for MWCP is straightforward: subclass the parser base, implement the run method to locate-decode-parse the sample it is handed, and report fields through the standard report object (report.add(...) with typed metadata such as a socket address, a mutex or an encryption key). MWCP then handles input, output formatting, and running your parser alongside others. The same locate-decode-parse body you write below drops into that method almost unchanged.

Robustness: writing for the next variant

The code you analysed is one build of a living family. The next sample will differ — a new key, a reordered field, an extra flag, a moved marker — and a brittle extractor silently returns wrong data, which is worse than returning nothing.

  • Fail loudly. If the marker is absent, the decode produces non-printable hosts, or a length runs past the buffer, raise a clear error and exit non-zero. A wrong config that looks plausible poisons every downstream system that trusts it.
  • Validate what you parse. Check that a host looks like a host, a port is in range, a length matches the bytes available. These checks also let you reject false marker hits — the same magic bytes can appear as a compare literal in code, not only as the real config.
  • Test against several samples. One sample is an anecdote. Collect a handful from the same family and run the extractor across all of them; the fields that vary are the real config, the fields that never change might be your parsing artefacts. Keep those samples as regression tests, exactly as the analysis pipeline keeps a regression corpus for rules.
  • Version your extractor. When you change it for a new variant, record which variant and keep the old behaviour testable, so you can tell whether a different result means the sample changed or your code did.
  • Do not run the sample to get the config unless you must. Static extraction is repeatable and safe; fall back to a debugger or emulation only when the config is fetched or decrypted with runtime-only state.

Warning: Obfuscation aimed at analysts is often aimed at extractors too. API hashing, control-flow flattening and per-build key derivation are meant to break a marker-based rule. When a YARA locate stops matching across builds, anchor the rule on the decoding code rather than the config data, which changes less often.

Sharing configurations responsibly

An extracted config is intelligence, and some of it is sensitive. C2 domains and mutex names are safe to share widely and belong in blocklists and feeds. But a config can also contain a victim identifier, an internal hostname baked into a targeted build, or infrastructure whose exposure tells an attacker they have been caught. Apply the same discipline the pipeline lesson applies to sample uploads: share by the incident's traffic-light protocol, prefer hash- and IOC-level sharing to posting whole samples, and defang URLs and hosts in any report so a config pasted into a ticket or a chat cannot be clicked. Platforms like MISP and MWDB exist to share configs with structure and access control; a public paste does not.

Lab: an extractor with a marker, a decoder and a loud failure

You will add a small encrypted config to a harmless sample, then write an extractor that locates the marker with YARA, decodes and parses the blob into JSON, and fails with a clear error on a variant where the marker is missing. Continue in the m8b directory from the previous lesson, or set up a fresh one with mingw-w64 and a venv holding pefile and yara-python.

  1. Extend gen.py to emit a config blob: an 8-byte ASCII marker in the clear, followed by the fields XORed with a one-byte key. The volatile qualifier stops -O2 from constant-folding the whole decode away, so the encoded bytes really appear in .data as they would in a real sample.

    python
    import struct
    MARKER, CFG_KEY = b"LABCFG01", 0x3D
    host, cid = b"update.example.com", b"lab-demo"
    body  = struct.pack("<B", len(host)) + host
    body += struct.pack("<H", 8443) + struct.pack("<H", 60)   # port, sleep
    body += struct.pack("<B", len(cid)) + cid
    blob = MARKER + bytes(b ^ CFG_KEY for b in body)          # marker left clear
    with open("lab_data.h", "a") as f:
        f.write("static volatile unsigned char CFG_BLOB[] = { %s };\n"
                % ", ".join("0x%02x" % b for b in blob))
        f.write("#define CFG_BLOB_LEN %d\n" % len(blob))
  2. Add the runtime decode to lab_sample.c so the built sample is self-consistent (it prints the config it carries):

    c
    /* in main(), after the strings: locate marker, XOR remainder with 0x3D */
    volatile unsigned char *p = CFG_BLOB;
    char marker[8];
    for (int i = 0; i < 8; i++) marker[i] = p[i];
    if (memcmp(marker, "LABCFG01", 8) == 0) {
        unsigned char cfg[64];
        for (int i = 0; i < CFG_BLOB_LEN - 8; i++) cfg[i] = p[8 + i] ^ 0x3D;
        int hlen = cfg[0];
        char host[64]; memcpy(host, cfg + 1, hlen); host[hlen] = '\0';
        unsigned port  = cfg[1 + hlen] | (cfg[2 + hlen] << 8);
        unsigned sleep = cfg[3 + hlen] | (cfg[4 + hlen] << 8);
        int ilen = cfg[5 + hlen];
        char id[64]; memcpy(id, cfg + 6 + hlen, ilen); id[ilen] = '\0';
        printf("config: host=%s port=%u sleep=%u id=%s\n", host, port, sleep, id);
    }

    Rebuild and confirm it prints the config:

    bash
    .venv/bin/python gen.py
    x86_64-w64-mingw32-gcc -O2 -s -o lab_sample.exe lab_sample.c
    text
    config: host=update.example.com port=8443 sleep=60 id=lab-demo
  3. Note a subtlety before writing the extractor: the marker string appears twice in the binary. Once as the memcmp comparison literal in .text (followed by code), and once as the real config blob in .data (followed by the XORed body). The extractor must try each hit and accept only the one that parses — a small lesson in why validation, not just location, is required.

    bash
    python3 -c "d=open('lab_sample.exe','rb').read(); i=0
    while (i:=d.find(b'LABCFG01', i)) >= 0: print(hex(i), d[i+8:i+13].hex()); i+=1"
    text
    0x2055 48394424380f       # .text: comparison literal, then code
    0x2200 2f484d595c         # .data: 0x2f = 0x12 ^ 0x3d, the real body
  4. Write config_extract.py on the locate-decode-parse-output pattern:

    python
    #!/usr/bin/env python3
    """config_extract.py - locate, decode and parse the config into JSON."""
    import json, struct, sys
    import yara
    
    CFG_KEY = 0x3D
    RULE = r'''
    rule lab_config_marker {
        strings: $marker = "LABCFG01"
        condition: $marker
    }'''
    
    class BadConfig(Exception):
        pass
    
    def parse_blob(data, off):
        """Decode (XOR) and parse one candidate; raise on anything implausible."""
        body = bytes(b ^ CFG_KEY for b in data[off + 8:])      # skip 8-byte marker
        try:
            host_len = body[0]
            host = body[1:1 + host_len]
            pos = 1 + host_len
            port, sleep = struct.unpack_from("<HH", body, pos)
            pos += 4
            id_len = body[pos]
            cid = body[pos + 1:pos + 1 + id_len]
        except (IndexError, struct.error) as e:
            raise BadConfig(f"structure does not unpack: {e}")
    
        if not (1 <= host_len <= 64 and len(host) == host_len):
            raise BadConfig(f"implausible host length {host_len}")
        if not all(0x20 <= c < 0x7f for c in host):
            raise BadConfig("host is not printable")
        if not (1 <= port <= 65535):
            raise BadConfig(f"port out of range: {port}")
        if not (1 <= id_len <= 64 and len(cid) == id_len):
            raise BadConfig("implausible campaign id")
    
        return {"c2_host": host.decode(), "c2_port": port,
                "sleep_seconds": sleep, "campaign_id": cid.decode(),
                "_source_offset": hex(off)}
    
    def extract(path):
        data = open(path, "rb").read()
        matches = yara.compile(source=RULE).match(data=data)
        hits = [inst.offset for m in matches for inst in m.strings[0].instances] \
            if matches else []
        if not hits:
            raise BadConfig("marker LABCFG01 not found: not this family, or a new variant")
        for off in hits:                       # try each hit; skip the code literal
            try:
                return parse_blob(data, off)
            except BadConfig:
                continue
        raise BadConfig(f"marker found at {[hex(h) for h in hits]} but none parsed")
    
    def main():
        if len(sys.argv) != 2:
            sys.exit(f"usage: {sys.argv[0]} sample.exe")
        try:
            print(json.dumps(extract(sys.argv[1]), indent=2))
        except BadConfig as e:
            print(f"error: {e}", file=sys.stderr)
            sys.exit(2)
    
    if __name__ == "__main__":
        main()
  5. Run it on the sample:

    bash
    .venv/bin/python config_extract.py lab_sample.exe
    json
    {
      "c2_host": "update.example.com",
      "c2_port": 8443,
      "sleep_seconds": 60,
      "campaign_id": "lab-demo",
      "_source_offset": "0x2200"
    }

    The extractor skipped the .text marker at 0x2055 — decoding there yields an implausible host length, so parse_blob rejected it — and parsed the real blob at 0x2200.

  6. Prove it fails loudly on a variant. Overwrite the marker to simulate a build that changed it, then re-run:

    bash
    python3 -c "d=bytearray(open('lab_sample.exe','rb').read()); i=0
    while (i:=d.find(b'LABCFG01', i)) >= 0: d[i:i+8]=b'ZZZZZZZZ'; i+=8
    open('lab_variant.exe','wb').write(d)"
    .venv/bin/python config_extract.py lab_variant.exe; echo "exit=$?"
    text
    error: marker LABCFG01 not found: not this family, or a new variant
    exit=2

    The non-zero exit and the clear message are the point: a pipeline running this extractor learns that the config was not found, rather than silently recording an empty or wrong config.

  7. Package it, conceptually, for MWCP. Move the extract body into a parser subclass's run method, report each field with its type (the C2 as a socket address, the id as a mission/campaign id, the interval as an integer), and let MWCP handle input and JSON output. Attach the YARA rule so the framework runs your parser only on samples that match. The locate-decode-parse logic does not change — only how the result is reported.

Questions to answer: Why does keying the YARA rule on the marker rather than on the decoding code make the extractor more fragile against a new build, and when would you choose each? Which validation check rejected the .text marker hit, and what would happen without it? If a second variant kept the marker but reordered the fields, which of the four stages would you edit? What single field in this config would you treat as sensitive and not put in a public blocklist, and why? How would you turn the three or four samples of a family into a regression test for this extractor?

Key takeaways

  • The configuration — C2 hosts and ports, keys, IDs, intervals, mutexes, paths, flags — is the highest-value output of an analysis, feeding IOCs, detections, remediation and intelligence at once.
  • Configs live in resources, overlays, .data blobs, .NET settings classes, the registry or a runtime fetch; high entropy in one of those is the tell.
  • Build every extractor as locate (YARA on a marker or the decode code), decode (reimplement the crypto), parse (typed fields) and output (normalised JSON), written as separate stages.
  • Reuse frameworks — MWCP, CAPE extractors, malduck, RATDecoders-style scripts — which all organise around a per-family extractor plus a runner; your code is just the family-specific part.
  • Make it robust: fail loudly, validate every parsed field, test against several samples and version the extractor, so a new build never yields a plausible wrong config.
  • Share configs by the incident's TLP, prefer IOC- and hash-level sharing and structured platforms to public pastes, and defang hosts in every report.