Lesson 8.1 · Encoding, Crypto & Signatures· 45 min
Recognising XOR, Base64 and Custom Encodings
Spot XOR, rolling-key and custom Base64 encodings in data and in code, recover their keys with key-length tests and known plaintext, and decode them.
Objectives
- Classify an unknown blob as hex, Base64, a custom alphabet, XOR or something stronger from its character set, length and entropy
- Recognise XOR, rolling-key, ADD/SUB/ROL and lookup-table decoding loops in optimised disassembly
- Recover a multi-byte XOR key with key-length tests and a known-plaintext crib
- Decode custom-alphabet Base64 and brute-force a rolling key in Python and CyberChef
- Avoid the classic traps: null-preserving XOR, key-length multiples and encoding layered on compression
An encoding changes how data looks without needing a secret that is hard
to obtain. Malware uses encodings everywhere a defender might otherwise read
plain text: in the configuration block that holds C2 hosts and campaign IDs, in
the strings the code decodes just before use, in files it stages on disk and in
the bytes it sends over the network. None of this is strong protection. It is
there to beat strings, naive signatures and a hurried analyst.
Strings and Obfuscated Strings showed how to notice that text is hidden, how FLOSS recovers decoded strings by emulation, and how to brute-force a single-byte XOR key. This lesson goes one step further: identifying which encoding you are looking at, from the data and from the code, and recovering the parameters when a 255-key brute force is no longer enough. The next lesson, Identifying Cryptographic Algorithms, takes over when the transformation is a real cipher.
Why an analyst cares
Decoding is rarely the goal. It is the step that unlocks the outputs your consumers need:
| Where the encoding sits | What decoding gives you | Who uses it |
|---|---|---|
| Embedded configuration | C2 hosts, ports, campaign and bot IDs, sleep intervals, mutex names | SOC (IOCs), threat intel (clustering) |
| Individual strings | API names, registry paths, commands, file names | You (capability assessment), detection engineers |
| Network traffic | Beacon format, command IDs, exfiltrated fields | Network detection (see Writing Network Signatures, later in this module) |
| Staged files | What was collected before exfiltration | Incident responders |
There is also a detection angle. Once you know a family's scheme and key, the encoded form of a known string is a perfectly good YARA anchor, and the decoding loop itself is often a more stable signature than any IOC.
Recognising encodings in data
Start with the blob itself. Four measurements settle most cases: the set of characters used, the length, the padding and the entropy.
Character set and length
| Encoding | Alphabet | Length rule | Other tells |
|---|---|---|---|
| Hex (Base16) | 0-9a-f or 0-9A-F | Even | Never mixes cases within one blob |
| Base32 | A-Z2-7 | Multiple of 8 when padded | = padding up to six characters |
| Base64 (RFC 4648) | A-Za-z0-9+/ | Multiple of 4 when padded | Zero, one or two = at the end only |
| Base64url | A-Za-z0-9-_ | Often unpadded | Common in URLs and tokens |
| Custom Base64 | Any 64 printable characters | As Base64 | Standard decoding "succeeds" but yields garbage |
| XOR, ADD, ROL, table substitution | Any byte value | Same as plaintext | Structure of the plaintext survives |
Base64 turns every 3 input bytes into 4 output characters, which is why the
padded length is a multiple of 4. Malware frequently strips the = padding to
look less like Base64. Before you decide a string is not Base64, add padding
until the length is a multiple of 4 and try again.
A custom alphabet is the more interesting case. Because the author only
permutes the 64-character table, the output still uses letters, digits and two
symbols, and a strict Base64 validator will happily accept it. The giveaway is
that the decoded bytes are not what you expect: random-looking binary where a
config should be. In the binary, look for the table itself: a 64-character
printable string with no repeats, often next to a lone = or the constant
0x3D in the encoder.
Entropy and byte frequency
Entropy is measured the same way as in Detecting Packers and Entropy, but here the ceiling tells you the most:
| Data | Typical entropy (bits per byte) | Why |
|---|---|---|
| English or config text | 4 to 5 | Few symbols, uneven frequencies |
| Hex text | at most 4.0 | Only 16 symbols are possible |
| Base64 text | at most 6.0 | Only 64 symbols (plus =) are possible |
| Single-byte XOR, ADD or table substitution of text | Same as the plaintext | The byte histogram is only relabelled |
| Repeating-key XOR of text | Somewhat higher than the plaintext | Each key byte relabels a different column |
| Compressed or encrypted data | 7.5 and above | Redundancy removed |
Two consequences are worth memorising. First, a simple substitution never changes the shape of the byte histogram. If a 4 KB blob has one byte value that dominates, it is probably an encoded run of zeros or spaces, and that byte is probably the key. Second, a Base64 blob whose decoded bytes have entropy near 8 is not "Base64-encoded data". It is Base64 wrapped around something compressed or encrypted, and you still have work to do.
Patterns that leak the key
XOR has a property that the author cannot switch off: 0x00 XOR k = k.
Wherever the plaintext contains zeros, the key appears in the clear.
- Single-byte XOR: runs of one repeated byte where zeros or padding were.
- Repeating-key XOR: the key itself, repeated, over the same regions. An encoded PE file shows the key over and over in its header and section padding, because PE files are full of zero bytes.
- Text encoded with a repeating key: equal plaintext characters in the same key column produce equal ciphertext bytes, so repeated substrings reappear at distances that are multiples of the key length.
The last point is the basis of key-length detection. XOR two ciphertext blocks that were encoded with the same key bytes and the key cancels out, leaving the XOR of two plaintext blocks. For text, that difference has few set bits. So the candidate key length with the lowest normalised Hamming distance between consecutive blocks is usually the right one, or a multiple of it.
Recognising encodings in code
In the disassembler, decoders share a shape: a short loop that reads a byte,
transforms it with one to three arithmetic instructions, and writes it back.
Recognising Compiler Idioms taught you to read
loops. Here is what the transformation inside them tends to look like, compiled
with GCC 15 at -O2 (alignment nops removed).
Repeating-key XOR. The index is reduced modulo the key length with and
when the length is a power of two; other lengths produce a div or a
magic-number multiplication:
xor4:
20: mov r9,rax
23: and r9d,0x3 ; i & 3 -> 4-byte key
27: movzx r9d,BYTE PTR [r8+r9*1] ; key[i & 3]
2c: xor BYTE PTR [rcx+rax*1],r9b
30: add rax,0x1
34: cmp rdx,rax
37: jne 20 <xor4+0x20>Rolling key. The key is updated on every iteration. Here the decoder subtracts a seed that grows by 3 each byte:
roll_sub:
50: sub BYTE PTR [rcx],r8b ; b[i] -= seed
53: add rcx,0x1
57: add r8d,0x3 ; seed += 3
5b: cmp rcx,rdx
5e: jne 50 <roll_sub+0x10>ADD/SUB and rotate combinations. ADD and SUB undo each other, and so do
rol and ror, so the decoder contains the inverse of whatever the encoder
did, in reverse order:
rol_add:
80: movzx eax,BYTE PTR [rcx]
83: add rcx,0x1
87: sub eax,0x21
8a: rol al,0x3
8d: mov BYTE PTR [rcx-0x1],alLookup-table substitution. Each byte indexes a 256-byte table. Custom
Base64 decoders use the same movzx pattern with a 64- or 256-entry reverse
table, plus shifts by 2, 4 and 6 and masks with 0x3f:
subst:
c0: movzx eax,BYTE PTR [rcx]
c3: add rcx,0x1
c7: movzx eax,BYTE PTR [r8+rax*1] ; table[b[i]]
cc: mov BYTE PTR [rcx-0x1],alA few rules keep you from chasing the wrong loops:
xor eax,eaxand other self-XORs only clear a register. Search for XOR with two different operands inside a loop. capa's "encode data using XOR" rule applies exactly this filter.- A chained (or "loopback") scheme uses the previous ciphertext or plaintext byte as the next key. Look for the value written in one iteration being read as the key in the next.
- A 32-bit immediate key such as
xor DWORD PTR [rax],0x37a21f4bis stored little-endian, so the key bytes are4b 1f a2 37. See endianness. - Decoders are usually called from many places with a pointer and a length.
A function with dozens of callers, each passing a different
.dataaddress, is the classic XOR string encryption layout, and the call sites tell you where every encoded string lives.
Recovering the key
When you have the decoder, reading the key out of it is the most reliable method. When you only have data, or want to confirm your reading, three techniques cover most simple schemes.
Known plaintext. You almost always know something about the plaintext:
an executable starts with MZ and contains This program cannot be run in DOS mode, a URL contains http, a config may contain = or ; at regular
places, and a family you have seen before uses known field names. For XOR,
ciphertext XOR plaintext = key stream. If the recovered stream repeats with
period n, you have the key and its length in one step.
Key-length detection. Without a crib at a known offset, score each candidate length. The Hamming-distance test above is the classic one. On short blobs a second test is more robust: split the data into n columns and check whether each column can be decoded to plausible characters by a single key byte. Only the right length (and its multiples) passes for every column.
Brute force. A single-byte key has 255 candidates, and a rolling key with a seed and a step has only 65,536. Score each candidate by printability or by a crib. That is milliseconds in Python.
Tip: When a key-length test likes both 4 and 8, the answer is 4. Any multiple of the true key length also aligns every column with a single key byte. Always take the smallest length that works, then confirm by decoding.
CyberChef recipes
CyberChef is quicker than code for a one-off and makes a good cross-check for your scripts:
| Goal | Recipe |
|---|---|
| Repeating-key XOR | From Hex, then XOR with the key as Hex (for example 4b1fa237) |
| Find a short XOR key | XOR Brute Force with a crib such as http or server= |
| Custom Base64 | From Base64 with the alphabet field replaced by the 64 characters from the binary |
| Translate an alphabet first | Substitute (custom to standard), then ordinary From Base64 |
| Base64 around compression | From Base64, then Zlib Inflate, Raw Inflate or Gunzip |
| "What is this?" | Magic with intensive mode, then check its guess by hand |
Rolling keys with a step, and anything with state between bytes, are awkward to express in CyberChef. Write ten lines of Python instead. The lab does exactly that, and the lesson Scripting String Decryption later in this module turns the approach into a reusable decoder.
Pitfalls
Null-preserving XOR. Some encoders skip bytes that are 0x00 or equal to
the key, so zeros stay zeros and the key never leaks. The decoder has the same
two comparisons in front of the xor. If you decode such data with a plain
XOR, every byte that should have been zero or the key comes out wrong, and your
crib search may fail on short strings. Mirror the exact skip condition.
Encoding over compression. Authors often compress first and encode second,
because compressed data is smaller and the encoding hides the compressor's
headers. After decoding, check the first bytes before assuming encryption:
78 9C or 78 DA is zlib, 1F 8B is gzip, PK is ZIP. A call to
RtlDecompressBuffer in the imports means LZNT1 or Xpress. High entropy after
a correct decode is expected for compressed data.
Layering in general. Real configs stack two or three steps: Base64 of XOR of zlib, or hex of RC4. Peel one layer, re-measure character set and entropy, and repeat. The order in the code is the reverse of the order you apply.
Look-alike hex and Base64. A Base64 string can consist only of hex characters by chance, and GUIDs, hashes and certificate thumbprints are legitimate hex. Context (who reads the value, and what the code does with it) decides, not the character set alone.
Lab: three encodings, three recoveries
You will generate three harmless encoded copies of a fake config and recover each one. Work in a scratch directory with a virtual environment:
python3 -m venv venv && . venv/bin/activate-
Generate the blobs. Save as
make_blobs.pyand run it:python # make_blobs.py - produce three harmless encoded copies of a demo config import base64 CONFIG = b"server=update.example.com;port=8080;id=lab" # 1. Repeating 4-byte XOR key XOR_KEY = bytes([0x4B, 0x1F, 0xA2, 0x37]) xor_blob = bytes(b ^ XOR_KEY[i % len(XOR_KEY)] for i, b in enumerate(CONFIG)) # 2. Base64 with a custom (shuffled) alphabet STD = b"ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/" CUSTOM = b"ZYXWVUTSRQPONMLKJIHGFEDCBAzyxwvutsrqponmlkjihgfedcba9876543210+/" b64_blob = base64.b64encode(CONFIG).translate(bytes.maketrans(STD, CUSTOM)) # 3. Rolling ADD: each byte gets (seed + 3*i) added, modulo 256 SEED = 0x5C roll_blob = bytes((b + SEED + 3 * i) & 0xFF for i, b in enumerate(CONFIG)) open("xor.bin", "wb").write(xor_blob) open("b64.txt", "wb").write(b64_blob) open("roll.bin", "wb").write(roll_blob) print("xor :", xor_blob.hex()) print("b64 :", b64_blob.decode()) print("roll:", roll_blob.hex())text xor : 387ad0412e6d9f423b7bc3432e31c74f2a72d25b2e31c1582624d258396b9f0f7b27920c227b9f5b2a7d b64 : x7EbwnEbKCEdATU9AH4ovTUgxTcoOnMeyGgdy6Q9KGtdLWZ2zDJ0yTUr roll: cfc4d4dbcdddabe6e4dbdbf1e5b1eb01edfc0201fdc9011011e21a1c2227f3f1ecf7f200312f0b3d3539From here on, pretend you did not write this script. You only have the three files and whatever the "sample" would tell you.
-
Recover the XOR key. Save as
xor_recover.py. It scores key lengths 1 to 8 with both tests, then applies the cribserver=:python # xor_recover.py - key length + known-plaintext recovery for repeating-key XOR import string, sys from itertools import combinations data = open(sys.argv[1], "rb").read() CRIB = sys.argv[2].encode() if len(sys.argv) > 2 else b"server=" ALLOWED = set((string.ascii_letters + string.digits + "=;.:/_-&?").encode()) def hamming(a: bytes, b: bytes) -> int: return sum(bin(x ^ y).count("1") for x, y in zip(a, b)) # 1. Key length, method A: normalised Hamming distance between blocks print("keylen hamming/bit columns-with-a-clean-key") for k in range(1, 9): blocks = [data[i:i + k] for i in range(0, len(data) - k + 1, k)][:4] pairs = list(combinations(blocks, 2)) ham = sum(hamming(a, b) for a, b in pairs) / len(pairs) / k # method B: can every column be decoded to "config-like" bytes by one key byte? clean = 0 for c in range(k): col = data[c::k] if any(all((x ^ kb) in ALLOWED for x in col) for kb in range(256)): clean += 1 print(f"{k:>6} {ham:>11.2f} {clean}/{k}") # 2. Known plaintext: ciphertext XOR crib = keystream stream = bytes(c ^ p for c, p in zip(data, CRIB)) print("\nkeystream under crib", CRIB, "=", stream.hex(" ")) for k in range(1, len(stream)): if all(stream[i] == stream[i % k] for i in range(len(stream))): key = stream[:k] break print("smallest period:", len(key), "key =", key.hex()) plain = bytes(b ^ key[i % len(key)] for i, b in enumerate(data)) print("decoded:", plain.decode(errors="replace"))bash python xor_recover.py xor.bintext keylen hamming/bit columns-with-a-clean-key 1 3.83 0/1 2 4.17 0/2 3 4.33 0/3 4 2.71 4/4 5 4.10 0/5 6 4.17 2/6 7 4.38 0/7 8 2.65 8/8 keystream under crib b'server=' = 4b 1f a2 37 4b 1f a2 smallest period: 4 key = 4b1fa237 decoded: server=update.example.com;port=8080;id=labBoth tests single out 4 and its multiple 8. Length 8 even has a slightly lower Hamming score, which is why you take the smallest length that works. The crib settles it: the key stream repeats with period 4.
-
Recognise and decode the custom Base64. Save as
b64_recover.py:python # b64_recover.py - spot a Base64-shaped blob, then decode it with a custom alphabet import base64, string, sys blob = open(sys.argv[1], "rb").read().strip() STD = b"ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/" # 1. Does it *look* like Base64? chars = set(blob.decode()) print("length", len(blob), "| length % 4 =", len(blob) % 4, "| distinct chars", len(chars), "| outside std alphabet:", sorted(chars - set(STD.decode() + "=")) or "none") # 2. Standard decoding "works" but produces garbage: a hint the alphabet is custom std = base64.b64decode(blob + b"=" * (-len(blob) % 4)) printable = sum(chr(b) in string.printable for b in std) / len(std) print(f"standard decode: {printable:.0%} printable -> {std[:24]!r}") # 3. Partial alphabet from a crib: each 3 plaintext bytes fix 4 index->char pairs crib = b"server=update" # expected start of the config std_enc = base64.b64encode(crib[: len(crib) // 3 * 3]) mapping = {chr(s): chr(c) for s, c in zip(std_enc, blob)} print("crib gives", len(mapping), "of 64 alphabet slots:", " ".join(f"{k}->{v}" for k, v in sorted(mapping.items()))) # 4. Full alphabet recovered from the binary (the 64-char string near the decoder) CUSTOM = b"ZYXWVUTSRQPONMLKJIHGFEDCBAzyxwvutsrqponmlkjihgfedcba9876543210+/" assert len(set(CUSTOM)) == 64 plain = base64.b64decode(blob.translate(bytes.maketrans(CUSTOM, STD)) + b"=" * (-len(blob) % 4)) print("decoded:", plain.decode())text length 56 | length % 4 = 0 | distinct chars 36 | outside std alphabet: none standard decode: 60% printable -> b'\xc7\xb1\x1b\xc2q\x1b(!\x1d\x015=\x00~(\xbd5 \xc57(:s\x1e' crib gives 13 of 64 alphabet slots: 0->9 2->7 F->U G->T P->K V->E X->C Z->A c->x d->w m->n w->d y->b decoded: server=update.example.com;port=8080;id=labThe blob passes every Base64 shape test, yet standard decoding yields binary junk. The crib recovers only 13 of the 64 slots, but they already show the pattern (
Z->A,X->C,V->E): each character range is reversed. In a real sample you would confirm this by finding the 64-character table in the binary. Then paste that table into CyberChef's From Base64 alphabet field and check you get the same output. -
Brute-force the rolling key. You have seen a decoder of the
roll_subshape (a seed that grows by a constant), but not its constants. Save asroll_brute.py:python # roll_brute.py - brute-force a rolling ADD key: out[i] = in[i] + seed + step*i import string, sys data = open(sys.argv[1], "rb").read() OK = set((string.ascii_letters + string.digits + "=;.:/_-").encode()) def decode(seed: int, step: int) -> bytes: return bytes((b - seed - step * i) & 0xFF for i, b in enumerate(data)) hits = [] for seed in range(256): for step in range(256): out = decode(seed, step) score = sum(c in OK for c in out) / len(out) if score == 1.0: hits.append((seed, step, out)) print(f"tried 65536 (seed, step) pairs, {len(hits)} fully printable") for seed, step, out in hits[:5]: print(f"seed=0x{seed:02x} step={step:<3} {out.decode()}") print("with crib 'server=':", [(hex(s), st) for s, st, o in hits if o.startswith(b"server=")])text tried 65536 (seed, step) pairs, 1 fully printable seed=0x5c step=3 server=update.example.com;port=8080;id=lab with crib 'server=': [('0x5c', 3)]The whole search takes about 0.3 seconds. Only one of 65,536 candidates decodes every byte to a config-like character.
-
Cross-check in CyberChef. Paste the XOR blob's hex, apply From Hex and XOR with key
4b1fa237(Hex). Then paste the Base64 blob, apply From Base64 and replace the alphabet with the custom table. Both should reproduce the config.
Questions to answer: Why did the Hamming test score length 8 slightly better than 4, and what rule protects you from that? The Base64 blob contains only 36 distinct characters. Would you still call it Base64 if it had been 12 characters long, and what else would you check? How would the XOR recovery change if the config had been compressed before encoding? If the rolling key's step depended on the previous plaintext byte, why would the 65,536-candidate brute force stop working?
Key takeaways
- Encodings hide configs, strings and traffic from casual inspection. Decoding them produces IOCs, protocol knowledge and detection content.
- Character set, length, padding and entropy classify most blobs: hex tops out at 4 bits per byte, Base64 at 6, and simple substitutions keep the plaintext's histogram shape.
- In code, decoders are short loops with a two-operand XOR, ADD/SUB, rotate or table lookup. The index arithmetic reveals the key length, and a changing key register reveals a rolling scheme.
- Known plaintext (
MZ,http, field names) turns XOR data into its key stream. Key-length tests also accept multiples of the true length, so take the smallest. - Watch for null-preserving XOR, stripped Base64 padding and compression under the encoding. Re-measure after every layer you peel.