Leçon 8.4 · Encodage, crypto & signatures· 55 min
Extracting Malware Configurations
Turn a sample's embedded configuration into normalised JSON with a locate-decode-parse-output extractor, and share the results responsibly.
Cette leçon n’est disponible qu’en anglais pour le moment.
Objectifs
- Explain what a malware configuration holds and why it is the highest-value output of an analysis
- Enumerate where configurations live: resources, overlays, .data blobs, .NET settings, the registry and runtime fetches
- Build an extractor on the locate–decode–parse–output pattern that emits normalised JSON
- Describe how MWCP, CAPE extractors and RATDecoders-style projects organise extractors, and package one for MWCP
- Make an extractor robust to version drift and share configurations responsibly
Of all the outputs an analysis produces, the configuration is the one the rest of the team wants first. A working triage report says what a sample is; the config says who it talks to and how. Extract it in an hour and you hand the SOC a set of live indicators and the detection engineers something to hunt on, long before a full reversing effort finishes. This lesson is about producing that config as clean, structured data — and about writing an extractor that keeps producing it as the family evolves.
It builds on the two skills either side of it in Module 8: recognising encodings and identifying crypto tell you how the config is protected, and scripting string decryption gives you the argument-recovery and reimplementation techniques you reuse here on a larger structure.
What a configuration holds
A malware family's config is the set of values that change from campaign to campaign while the code stays the same. Pulling them out tells you almost everything operationally important:
| Field | Why it matters |
|---|---|
| C2 hosts and ports | Network IOCs; the core of any blocklist or hunt |
| Fallback / backup C2 lists | Infrastructure you would otherwise miss |
| Campaign or botnet ID | Clusters samples into campaigns; attribution pivot |
| Encryption keys / RC4 keys | Decrypt the C2 traffic and later payloads |
| Sleep / beacon interval | Behavioural signature; tunes detection windows |
| Mutex names | Host IOC; a reliable, cheap detection |
| Install paths and filenames | Remediation; where to clean |
| Registry keys | Persistence locations to remove |
| Feature flags | What the sample is configured to actually do |
| Version string | Tracks the family's development over time |
Read that list against the output table from What Is Malware Analysis?: a single extracted config feeds IOCs, detection content, remediation guidance and intelligence at once. That is why config extraction, not full reversing, is where most analytical value per hour is created.
Where configurations live
Before you can decode a config you have to find it. Families hide it in a handful of predictable places:
- PE resources. The
.rsrcdirectory is a favourite: a config sits in a custom-named resource, often encrypted, sometimes disguised as an icon or bitmap. The resources and overlays lesson covers how to enumerate them. - Overlays. Data appended after the last section, past the end of the PE as the headers describe it, is invisible to naive parsers and a common config hiding place.
.data/.rdatablobs. A byte array in a data section, decrypted at startup. Usually preceded or wrapped by a marker — a magic string or byte sequence the code searches for. Markers are the single most useful thing an extractor can key on.- .NET settings classes. As covered in Analysing .NET
Malware, managed families concentrate the config in
a static settings class assigned in the
.cctor, often Base64 then AES. - The registry or a dropped file. Some families write their config to the host at install and read it back, so the config is not in the original sample at all — you recover it from the infected host or a memory dump.
- Fetched at runtime. The worst case: the sample downloads its config from C2 on first run. Static extraction gets you the decryption key and the URL; the config itself needs a live or simulated C2.
High entropy in a resource, an overlay or a .data
blob is the tell that one of these holds an encrypted config.
The extractor pattern
Every config extractor, however different the family, has the same four stages. Keeping them separate is what makes an extractor readable, testable and maintainable.
- Locate. Find the config bytes. A YARA rule on the marker or on the decoding code is the standard mechanism, because it survives minor rebuilds and gives you the offset. Failing a marker, locate the resource by name, the overlay by parsing the PE, or the blob by an xref from the decode routine.
- Decode. Reverse the protection. This is where the encoding and crypto skills from earlier in the module pay off: reimplement the XOR, RC4, AES or custom scheme exactly, reusing the key you recovered from the code.
- Parse. Impose structure on the decoded bytes. Configs are typed: a length
byte then a host, a fixed-offset table of fields, a length-prefixed list, a
serialized blob (Protocol Buffers, JSON, a .NET
BinaryFormatterstream). Read each field with its real type. - Output. Emit normalised data — JSON with stable field names — so every sample of the family produces the same shape regardless of which fields are present. Normalisation is what lets downstream tooling consume configs from many families uniformly.
Tip: Write the four stages as four functions even for a one-off. When the next variant changes only the key, you edit
decodeand nothing else; when it adds a field, you editparse. An extractor that mixes locating and parsing in one loop has to be rewritten for every variant.
Existing frameworks
You rarely start from a blank file. Several open frameworks give you the plumbing — sample loading, crypto helpers, a config schema, a runner — so your code is just the family-specific locate, decode and parse.
| Project | Maintainer | Shape | Where it fits |
|---|---|---|---|
| DC3-MWCP | DoD Cyber Crime Center | Framework where each parser is a Python class emitting a standard report of typed fields | Standardised extractor output across many families |
| CAPE extractors | Open-source (kevoreilly) | Per-family Python modules run automatically against dumped payloads in the CAPE sandbox | Extraction integrated with unpacking and detonation |
| malduck | CERT Polska | Library of crypto, PE and structure helpers, plus a YARA-driven extractor model | Building blocks and a modern extractor runner |
| RATDecoders-style projects | Community (e.g. StormShield, Malwoverview lineage) | A directory of standalone per-family scripts, each self-contained | Reading how a specific family is parsed; quick reference |
They are organised the same way underneath: a registry of per-family
extractors, each declaring how to recognise its family (usually a YARA rule)
and how to produce a config, plus a runner that matches a sample to an
extractor and normalises the result. MWCP calls the unit a parser; CAPE calls
it an extractor module; malduck calls it an extractor class with an embedded
yara_rule. Learning one transfers to the others.
Packaging your extractor for MWCP is straightforward: subclass the parser base,
implement the run method to locate-decode-parse the sample it is handed, and
report fields through the standard report object (report.add(...) with typed
metadata such as a socket address, a mutex or an encryption key). MWCP then
handles input, output formatting, and running your parser alongside others. The
same locate-decode-parse body you write below drops into that method almost
unchanged.
Robustness: writing for the next variant
The code you analysed is one build of a living family. The next sample will differ — a new key, a reordered field, an extra flag, a moved marker — and a brittle extractor silently returns wrong data, which is worse than returning nothing.
- Fail loudly. If the marker is absent, the decode produces non-printable hosts, or a length runs past the buffer, raise a clear error and exit non-zero. A wrong config that looks plausible poisons every downstream system that trusts it.
- Validate what you parse. Check that a host looks like a host, a port is in range, a length matches the bytes available. These checks also let you reject false marker hits — the same magic bytes can appear as a compare literal in code, not only as the real config.
- Test against several samples. One sample is an anecdote. Collect a handful from the same family and run the extractor across all of them; the fields that vary are the real config, the fields that never change might be your parsing artefacts. Keep those samples as regression tests, exactly as the analysis pipeline keeps a regression corpus for rules.
- Version your extractor. When you change it for a new variant, record which variant and keep the old behaviour testable, so you can tell whether a different result means the sample changed or your code did.
- Do not run the sample to get the config unless you must. Static extraction is repeatable and safe; fall back to a debugger or emulation only when the config is fetched or decrypted with runtime-only state.
Warning: Obfuscation aimed at analysts is often aimed at extractors too. API hashing, control-flow flattening and per-build key derivation are meant to break a marker-based rule. When a YARA locate stops matching across builds, anchor the rule on the decoding code rather than the config data, which changes less often.
Sharing configurations responsibly
An extracted config is intelligence, and some of it is sensitive. C2 domains and mutex names are safe to share widely and belong in blocklists and feeds. But a config can also contain a victim identifier, an internal hostname baked into a targeted build, or infrastructure whose exposure tells an attacker they have been caught. Apply the same discipline the pipeline lesson applies to sample uploads: share by the incident's traffic-light protocol, prefer hash- and IOC-level sharing to posting whole samples, and defang URLs and hosts in any report so a config pasted into a ticket or a chat cannot be clicked. Platforms like MISP and MWDB exist to share configs with structure and access control; a public paste does not.
Lab: an extractor with a marker, a decoder and a loud failure
You will add a small encrypted config to a harmless sample, then write an
extractor that locates the marker with YARA, decodes and parses the blob into
JSON, and fails with a clear error on a variant where the marker is missing.
Continue in the m8b directory from the previous lesson, or set up a fresh one
with mingw-w64 and a venv holding pefile and yara-python.
-
Extend
gen.pyto emit a config blob: an 8-byte ASCII marker in the clear, followed by the fields XORed with a one-byte key. Thevolatilequalifier stops-O2from constant-folding the whole decode away, so the encoded bytes really appear in.dataas they would in a real sample.python import struct MARKER, CFG_KEY = b"LABCFG01", 0x3D host, cid = b"update.example.com", b"lab-demo" body = struct.pack("<B", len(host)) + host body += struct.pack("<H", 8443) + struct.pack("<H", 60) # port, sleep body += struct.pack("<B", len(cid)) + cid blob = MARKER + bytes(b ^ CFG_KEY for b in body) # marker left clear with open("lab_data.h", "a") as f: f.write("static volatile unsigned char CFG_BLOB[] = { %s };\n" % ", ".join("0x%02x" % b for b in blob)) f.write("#define CFG_BLOB_LEN %d\n" % len(blob)) -
Add the runtime decode to
lab_sample.cso the built sample is self-consistent (it prints the config it carries):c /* in main(), after the strings: locate marker, XOR remainder with 0x3D */ volatile unsigned char *p = CFG_BLOB; char marker[8]; for (int i = 0; i < 8; i++) marker[i] = p[i]; if (memcmp(marker, "LABCFG01", 8) == 0) { unsigned char cfg[64]; for (int i = 0; i < CFG_BLOB_LEN - 8; i++) cfg[i] = p[8 + i] ^ 0x3D; int hlen = cfg[0]; char host[64]; memcpy(host, cfg + 1, hlen); host[hlen] = '\0'; unsigned port = cfg[1 + hlen] | (cfg[2 + hlen] << 8); unsigned sleep = cfg[3 + hlen] | (cfg[4 + hlen] << 8); int ilen = cfg[5 + hlen]; char id[64]; memcpy(id, cfg + 6 + hlen, ilen); id[ilen] = '\0'; printf("config: host=%s port=%u sleep=%u id=%s\n", host, port, sleep, id); }Rebuild and confirm it prints the config:
bash .venv/bin/python gen.py x86_64-w64-mingw32-gcc -O2 -s -o lab_sample.exe lab_sample.ctext config: host=update.example.com port=8443 sleep=60 id=lab-demo -
Note a subtlety before writing the extractor: the marker string appears twice in the binary. Once as the
memcmpcomparison literal in.text(followed by code), and once as the real config blob in.data(followed by the XORed body). The extractor must try each hit and accept only the one that parses — a small lesson in why validation, not just location, is required.bash python3 -c "d=open('lab_sample.exe','rb').read(); i=0 while (i:=d.find(b'LABCFG01', i)) >= 0: print(hex(i), d[i+8:i+13].hex()); i+=1"text 0x2055 48394424380f # .text: comparison literal, then code 0x2200 2f484d595c # .data: 0x2f = 0x12 ^ 0x3d, the real body -
Write
config_extract.pyon the locate-decode-parse-output pattern:python #!/usr/bin/env python3 """config_extract.py - locate, decode and parse the config into JSON.""" import json, struct, sys import yara CFG_KEY = 0x3D RULE = r''' rule lab_config_marker { strings: $marker = "LABCFG01" condition: $marker }''' class BadConfig(Exception): pass def parse_blob(data, off): """Decode (XOR) and parse one candidate; raise on anything implausible.""" body = bytes(b ^ CFG_KEY for b in data[off + 8:]) # skip 8-byte marker try: host_len = body[0] host = body[1:1 + host_len] pos = 1 + host_len port, sleep = struct.unpack_from("<HH", body, pos) pos += 4 id_len = body[pos] cid = body[pos + 1:pos + 1 + id_len] except (IndexError, struct.error) as e: raise BadConfig(f"structure does not unpack: {e}") if not (1 <= host_len <= 64 and len(host) == host_len): raise BadConfig(f"implausible host length {host_len}") if not all(0x20 <= c < 0x7f for c in host): raise BadConfig("host is not printable") if not (1 <= port <= 65535): raise BadConfig(f"port out of range: {port}") if not (1 <= id_len <= 64 and len(cid) == id_len): raise BadConfig("implausible campaign id") return {"c2_host": host.decode(), "c2_port": port, "sleep_seconds": sleep, "campaign_id": cid.decode(), "_source_offset": hex(off)} def extract(path): data = open(path, "rb").read() matches = yara.compile(source=RULE).match(data=data) hits = [inst.offset for m in matches for inst in m.strings[0].instances] \ if matches else [] if not hits: raise BadConfig("marker LABCFG01 not found: not this family, or a new variant") for off in hits: # try each hit; skip the code literal try: return parse_blob(data, off) except BadConfig: continue raise BadConfig(f"marker found at {[hex(h) for h in hits]} but none parsed") def main(): if len(sys.argv) != 2: sys.exit(f"usage: {sys.argv[0]} sample.exe") try: print(json.dumps(extract(sys.argv[1]), indent=2)) except BadConfig as e: print(f"error: {e}", file=sys.stderr) sys.exit(2) if __name__ == "__main__": main() -
Run it on the sample:
bash .venv/bin/python config_extract.py lab_sample.exejson { "c2_host": "update.example.com", "c2_port": 8443, "sleep_seconds": 60, "campaign_id": "lab-demo", "_source_offset": "0x2200" }The extractor skipped the
.textmarker at0x2055— decoding there yields an implausible host length, soparse_blobrejected it — and parsed the real blob at0x2200. -
Prove it fails loudly on a variant. Overwrite the marker to simulate a build that changed it, then re-run:
bash python3 -c "d=bytearray(open('lab_sample.exe','rb').read()); i=0 while (i:=d.find(b'LABCFG01', i)) >= 0: d[i:i+8]=b'ZZZZZZZZ'; i+=8 open('lab_variant.exe','wb').write(d)" .venv/bin/python config_extract.py lab_variant.exe; echo "exit=$?"text error: marker LABCFG01 not found: not this family, or a new variant exit=2The non-zero exit and the clear message are the point: a pipeline running this extractor learns that the config was not found, rather than silently recording an empty or wrong config.
-
Package it, conceptually, for MWCP. Move the
extractbody into a parser subclass'srunmethod, report each field with its type (the C2 as a socket address, the id as a mission/campaign id, the interval as an integer), and let MWCP handle input and JSON output. Attach the YARA rule so the framework runs your parser only on samples that match. The locate-decode-parse logic does not change — only how the result is reported.
Questions to answer: Why does keying the YARA rule on the marker rather than
on the decoding code make the extractor more fragile against a new build, and
when would you choose each? Which validation check rejected the .text marker
hit, and what would happen without it? If a second variant kept the marker but
reordered the fields, which of the four stages would you edit? What single field
in this config would you treat as sensitive and not put in a public blocklist,
and why? How would you turn the three or four samples of a family into a
regression test for this extractor?
Key takeaways
- The configuration — C2 hosts and ports, keys, IDs, intervals, mutexes, paths, flags — is the highest-value output of an analysis, feeding IOCs, detections, remediation and intelligence at once.
- Configs live in resources, overlays,
.datablobs, .NET settings classes, the registry or a runtime fetch; high entropy in one of those is the tell. - Build every extractor as locate (YARA on a marker or the decode code), decode (reimplement the crypto), parse (typed fields) and output (normalised JSON), written as separate stages.
- Reuse frameworks — MWCP, CAPE extractors, malduck, RATDecoders-style scripts — which all organise around a per-family extractor plus a runner; your code is just the family-specific part.
- Make it robust: fail loudly, validate every parsed field, test against several samples and version the extractor, so a new build never yields a plausible wrong config.
- Share configs by the incident's TLP, prefer IOC- and hash-level sharing and structured platforms to public pastes, and defang hosts in every report.