Skip to content

Lesson 8.1 · Encoding, Crypto & Signatures· 45 min

Recognising XOR, Base64 and Custom Encodings

Spot XOR, rolling-key and custom Base64 encodings in data and in code, recover their keys with key-length tests and known plaintext, and decode them.

Objectives

  • Classify an unknown blob as hex, Base64, a custom alphabet, XOR or something stronger from its character set, length and entropy
  • Recognise XOR, rolling-key, ADD/SUB/ROL and lookup-table decoding loops in optimised disassembly
  • Recover a multi-byte XOR key with key-length tests and a known-plaintext crib
  • Decode custom-alphabet Base64 and brute-force a rolling key in Python and CyberChef
  • Avoid the classic traps: null-preserving XOR, key-length multiples and encoding layered on compression

An encoding changes how data looks without needing a secret that is hard to obtain. Malware uses encodings everywhere a defender might otherwise read plain text: in the configuration block that holds C2 hosts and campaign IDs, in the strings the code decodes just before use, in files it stages on disk and in the bytes it sends over the network. None of this is strong protection. It is there to beat strings, naive signatures and a hurried analyst.

Strings and Obfuscated Strings showed how to notice that text is hidden, how FLOSS recovers decoded strings by emulation, and how to brute-force a single-byte XOR key. This lesson goes one step further: identifying which encoding you are looking at, from the data and from the code, and recovering the parameters when a 255-key brute force is no longer enough. The next lesson, Identifying Cryptographic Algorithms, takes over when the transformation is a real cipher.

Why an analyst cares

Decoding is rarely the goal. It is the step that unlocks the outputs your consumers need:

Where the encoding sitsWhat decoding gives youWho uses it
Embedded configurationC2 hosts, ports, campaign and bot IDs, sleep intervals, mutex namesSOC (IOCs), threat intel (clustering)
Individual stringsAPI names, registry paths, commands, file namesYou (capability assessment), detection engineers
Network trafficBeacon format, command IDs, exfiltrated fieldsNetwork detection (see Writing Network Signatures, later in this module)
Staged filesWhat was collected before exfiltrationIncident responders

There is also a detection angle. Once you know a family's scheme and key, the encoded form of a known string is a perfectly good YARA anchor, and the decoding loop itself is often a more stable signature than any IOC.

Recognising encodings in data

Start with the blob itself. Four measurements settle most cases: the set of characters used, the length, the padding and the entropy.

Character set and length

EncodingAlphabetLength ruleOther tells
Hex (Base16)0-9a-f or 0-9A-FEvenNever mixes cases within one blob
Base32A-Z2-7Multiple of 8 when padded= padding up to six characters
Base64 (RFC 4648)A-Za-z0-9+/Multiple of 4 when paddedZero, one or two = at the end only
Base64urlA-Za-z0-9-_Often unpaddedCommon in URLs and tokens
Custom Base64Any 64 printable charactersAs Base64Standard decoding "succeeds" but yields garbage
XOR, ADD, ROL, table substitutionAny byte valueSame as plaintextStructure of the plaintext survives

Base64 turns every 3 input bytes into 4 output characters, which is why the padded length is a multiple of 4. Malware frequently strips the = padding to look less like Base64. Before you decide a string is not Base64, add padding until the length is a multiple of 4 and try again.

A custom alphabet is the more interesting case. Because the author only permutes the 64-character table, the output still uses letters, digits and two symbols, and a strict Base64 validator will happily accept it. The giveaway is that the decoded bytes are not what you expect: random-looking binary where a config should be. In the binary, look for the table itself: a 64-character printable string with no repeats, often next to a lone = or the constant 0x3D in the encoder.

Entropy and byte frequency

Entropy is measured the same way as in Detecting Packers and Entropy, but here the ceiling tells you the most:

DataTypical entropy (bits per byte)Why
English or config text4 to 5Few symbols, uneven frequencies
Hex textat most 4.0Only 16 symbols are possible
Base64 textat most 6.0Only 64 symbols (plus =) are possible
Single-byte XOR, ADD or table substitution of textSame as the plaintextThe byte histogram is only relabelled
Repeating-key XOR of textSomewhat higher than the plaintextEach key byte relabels a different column
Compressed or encrypted data7.5 and aboveRedundancy removed

Two consequences are worth memorising. First, a simple substitution never changes the shape of the byte histogram. If a 4 KB blob has one byte value that dominates, it is probably an encoded run of zeros or spaces, and that byte is probably the key. Second, a Base64 blob whose decoded bytes have entropy near 8 is not "Base64-encoded data". It is Base64 wrapped around something compressed or encrypted, and you still have work to do.

Patterns that leak the key

XOR has a property that the author cannot switch off: 0x00 XOR k = k. Wherever the plaintext contains zeros, the key appears in the clear.

  • Single-byte XOR: runs of one repeated byte where zeros or padding were.
  • Repeating-key XOR: the key itself, repeated, over the same regions. An encoded PE file shows the key over and over in its header and section padding, because PE files are full of zero bytes.
  • Text encoded with a repeating key: equal plaintext characters in the same key column produce equal ciphertext bytes, so repeated substrings reappear at distances that are multiples of the key length.

The last point is the basis of key-length detection. XOR two ciphertext blocks that were encoded with the same key bytes and the key cancels out, leaving the XOR of two plaintext blocks. For text, that difference has few set bits. So the candidate key length with the lowest normalised Hamming distance between consecutive blocks is usually the right one, or a multiple of it.

Recognising encodings in code

In the disassembler, decoders share a shape: a short loop that reads a byte, transforms it with one to three arithmetic instructions, and writes it back. Recognising Compiler Idioms taught you to read loops. Here is what the transformation inside them tends to look like, compiled with GCC 15 at -O2 (alignment nops removed).

Repeating-key XOR. The index is reduced modulo the key length with and when the length is a power of two; other lengths produce a div or a magic-number multiplication:

asm
xor4:
  20:  mov    r9,rax
  23:  and    r9d,0x3                  ; i & 3  -> 4-byte key
  27:  movzx  r9d,BYTE PTR [r8+r9*1]   ; key[i & 3]
  2c:  xor    BYTE PTR [rcx+rax*1],r9b
  30:  add    rax,0x1
  34:  cmp    rdx,rax
  37:  jne    20 <xor4+0x20>

Rolling key. The key is updated on every iteration. Here the decoder subtracts a seed that grows by 3 each byte:

asm
roll_sub:
  50:  sub    BYTE PTR [rcx],r8b       ; b[i] -= seed
  53:  add    rcx,0x1
  57:  add    r8d,0x3                  ; seed += 3
  5b:  cmp    rcx,rdx
  5e:  jne    50 <roll_sub+0x10>

ADD/SUB and rotate combinations. ADD and SUB undo each other, and so do rol and ror, so the decoder contains the inverse of whatever the encoder did, in reverse order:

asm
rol_add:
  80:  movzx  eax,BYTE PTR [rcx]
  83:  add    rcx,0x1
  87:  sub    eax,0x21
  8a:  rol    al,0x3
  8d:  mov    BYTE PTR [rcx-0x1],al

Lookup-table substitution. Each byte indexes a 256-byte table. Custom Base64 decoders use the same movzx pattern with a 64- or 256-entry reverse table, plus shifts by 2, 4 and 6 and masks with 0x3f:

asm
subst:
  c0:  movzx  eax,BYTE PTR [rcx]
  c3:  add    rcx,0x1
  c7:  movzx  eax,BYTE PTR [r8+rax*1]  ; table[b[i]]
  cc:  mov    BYTE PTR [rcx-0x1],al

A few rules keep you from chasing the wrong loops:

  • xor eax,eax and other self-XORs only clear a register. Search for XOR with two different operands inside a loop. capa's "encode data using XOR" rule applies exactly this filter.
  • A chained (or "loopback") scheme uses the previous ciphertext or plaintext byte as the next key. Look for the value written in one iteration being read as the key in the next.
  • A 32-bit immediate key such as xor DWORD PTR [rax],0x37a21f4b is stored little-endian, so the key bytes are 4b 1f a2 37. See endianness.
  • Decoders are usually called from many places with a pointer and a length. A function with dozens of callers, each passing a different .data address, is the classic XOR string encryption layout, and the call sites tell you where every encoded string lives.

Recovering the key

When you have the decoder, reading the key out of it is the most reliable method. When you only have data, or want to confirm your reading, three techniques cover most simple schemes.

Known plaintext. You almost always know something about the plaintext: an executable starts with MZ and contains This program cannot be run in DOS mode, a URL contains http, a config may contain = or ; at regular places, and a family you have seen before uses known field names. For XOR, ciphertext XOR plaintext = key stream. If the recovered stream repeats with period n, you have the key and its length in one step.

Key-length detection. Without a crib at a known offset, score each candidate length. The Hamming-distance test above is the classic one. On short blobs a second test is more robust: split the data into n columns and check whether each column can be decoded to plausible characters by a single key byte. Only the right length (and its multiples) passes for every column.

Brute force. A single-byte key has 255 candidates, and a rolling key with a seed and a step has only 65,536. Score each candidate by printability or by a crib. That is milliseconds in Python.

Tip: When a key-length test likes both 4 and 8, the answer is 4. Any multiple of the true key length also aligns every column with a single key byte. Always take the smallest length that works, then confirm by decoding.

CyberChef recipes

CyberChef is quicker than code for a one-off and makes a good cross-check for your scripts:

GoalRecipe
Repeating-key XORFrom Hex, then XOR with the key as Hex (for example 4b1fa237)
Find a short XOR keyXOR Brute Force with a crib such as http or server=
Custom Base64From Base64 with the alphabet field replaced by the 64 characters from the binary
Translate an alphabet firstSubstitute (custom to standard), then ordinary From Base64
Base64 around compressionFrom Base64, then Zlib Inflate, Raw Inflate or Gunzip
"What is this?"Magic with intensive mode, then check its guess by hand

Rolling keys with a step, and anything with state between bytes, are awkward to express in CyberChef. Write ten lines of Python instead. The lab does exactly that, and the lesson Scripting String Decryption later in this module turns the approach into a reusable decoder.

Pitfalls

Null-preserving XOR. Some encoders skip bytes that are 0x00 or equal to the key, so zeros stay zeros and the key never leaks. The decoder has the same two comparisons in front of the xor. If you decode such data with a plain XOR, every byte that should have been zero or the key comes out wrong, and your crib search may fail on short strings. Mirror the exact skip condition.

Encoding over compression. Authors often compress first and encode second, because compressed data is smaller and the encoding hides the compressor's headers. After decoding, check the first bytes before assuming encryption: 78 9C or 78 DA is zlib, 1F 8B is gzip, PK is ZIP. A call to RtlDecompressBuffer in the imports means LZNT1 or Xpress. High entropy after a correct decode is expected for compressed data.

Layering in general. Real configs stack two or three steps: Base64 of XOR of zlib, or hex of RC4. Peel one layer, re-measure character set and entropy, and repeat. The order in the code is the reverse of the order you apply.

Look-alike hex and Base64. A Base64 string can consist only of hex characters by chance, and GUIDs, hashes and certificate thumbprints are legitimate hex. Context (who reads the value, and what the code does with it) decides, not the character set alone.

Lab: three encodings, three recoveries

You will generate three harmless encoded copies of a fake config and recover each one. Work in a scratch directory with a virtual environment:

bash
python3 -m venv venv && . venv/bin/activate
  1. Generate the blobs. Save as make_blobs.py and run it:

    python
    # make_blobs.py - produce three harmless encoded copies of a demo config
    import base64
    
    CONFIG = b"server=update.example.com;port=8080;id=lab"
    
    # 1. Repeating 4-byte XOR key
    XOR_KEY = bytes([0x4B, 0x1F, 0xA2, 0x37])
    xor_blob = bytes(b ^ XOR_KEY[i % len(XOR_KEY)] for i, b in enumerate(CONFIG))
    
    # 2. Base64 with a custom (shuffled) alphabet
    STD = b"ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/"
    CUSTOM = b"ZYXWVUTSRQPONMLKJIHGFEDCBAzyxwvutsrqponmlkjihgfedcba9876543210+/"
    b64_blob = base64.b64encode(CONFIG).translate(bytes.maketrans(STD, CUSTOM))
    
    # 3. Rolling ADD: each byte gets (seed + 3*i) added, modulo 256
    SEED = 0x5C
    roll_blob = bytes((b + SEED + 3 * i) & 0xFF for i, b in enumerate(CONFIG))
    
    open("xor.bin", "wb").write(xor_blob)
    open("b64.txt", "wb").write(b64_blob)
    open("roll.bin", "wb").write(roll_blob)
    print("xor :", xor_blob.hex())
    print("b64 :", b64_blob.decode())
    print("roll:", roll_blob.hex())
    text
    xor : 387ad0412e6d9f423b7bc3432e31c74f2a72d25b2e31c1582624d258396b9f0f7b27920c227b9f5b2a7d
    b64 : x7EbwnEbKCEdATU9AH4ovTUgxTcoOnMeyGgdy6Q9KGtdLWZ2zDJ0yTUr
    roll: cfc4d4dbcdddabe6e4dbdbf1e5b1eb01edfc0201fdc9011011e21a1c2227f3f1ecf7f200312f0b3d3539

    From here on, pretend you did not write this script. You only have the three files and whatever the "sample" would tell you.

  2. Recover the XOR key. Save as xor_recover.py. It scores key lengths 1 to 8 with both tests, then applies the crib server=:

    python
    # xor_recover.py - key length + known-plaintext recovery for repeating-key XOR
    import string, sys
    from itertools import combinations
    
    data = open(sys.argv[1], "rb").read()
    CRIB = sys.argv[2].encode() if len(sys.argv) > 2 else b"server="
    ALLOWED = set((string.ascii_letters + string.digits + "=;.:/_-&?").encode())
    
    def hamming(a: bytes, b: bytes) -> int:
        return sum(bin(x ^ y).count("1") for x, y in zip(a, b))
    
    # 1. Key length, method A: normalised Hamming distance between blocks
    print("keylen  hamming/bit  columns-with-a-clean-key")
    for k in range(1, 9):
        blocks = [data[i:i + k] for i in range(0, len(data) - k + 1, k)][:4]
        pairs = list(combinations(blocks, 2))
        ham = sum(hamming(a, b) for a, b in pairs) / len(pairs) / k
        # method B: can every column be decoded to "config-like" bytes by one key byte?
        clean = 0
        for c in range(k):
            col = data[c::k]
            if any(all((x ^ kb) in ALLOWED for x in col) for kb in range(256)):
                clean += 1
        print(f"{k:>6}  {ham:>11.2f}  {clean}/{k}")
    
    # 2. Known plaintext: ciphertext XOR crib = keystream
    stream = bytes(c ^ p for c, p in zip(data, CRIB))
    print("\nkeystream under crib", CRIB, "=", stream.hex(" "))
    for k in range(1, len(stream)):
        if all(stream[i] == stream[i % k] for i in range(len(stream))):
            key = stream[:k]
            break
    print("smallest period:", len(key), "key =", key.hex())
    plain = bytes(b ^ key[i % len(key)] for i, b in enumerate(data))
    print("decoded:", plain.decode(errors="replace"))
    bash
    python xor_recover.py xor.bin
    text
    keylen  hamming/bit  columns-with-a-clean-key
         1         3.83  0/1
         2         4.17  0/2
         3         4.33  0/3
         4         2.71  4/4
         5         4.10  0/5
         6         4.17  2/6
         7         4.38  0/7
         8         2.65  8/8
    
    keystream under crib b'server=' = 4b 1f a2 37 4b 1f a2
    smallest period: 4 key = 4b1fa237
    decoded: server=update.example.com;port=8080;id=lab

    Both tests single out 4 and its multiple 8. Length 8 even has a slightly lower Hamming score, which is why you take the smallest length that works. The crib settles it: the key stream repeats with period 4.

  3. Recognise and decode the custom Base64. Save as b64_recover.py:

    python
    # b64_recover.py - spot a Base64-shaped blob, then decode it with a custom alphabet
    import base64, string, sys
    
    blob = open(sys.argv[1], "rb").read().strip()
    STD = b"ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/"
    
    # 1. Does it *look* like Base64?
    chars = set(blob.decode())
    print("length", len(blob), "| length % 4 =", len(blob) % 4,
          "| distinct chars", len(chars),
          "| outside std alphabet:", sorted(chars - set(STD.decode() + "=")) or "none")
    
    # 2. Standard decoding "works" but produces garbage: a hint the alphabet is custom
    std = base64.b64decode(blob + b"=" * (-len(blob) % 4))
    printable = sum(chr(b) in string.printable for b in std) / len(std)
    print(f"standard decode: {printable:.0%} printable -> {std[:24]!r}")
    
    # 3. Partial alphabet from a crib: each 3 plaintext bytes fix 4 index->char pairs
    crib = b"server=update"            # expected start of the config
    std_enc = base64.b64encode(crib[: len(crib) // 3 * 3])
    mapping = {chr(s): chr(c) for s, c in zip(std_enc, blob)}
    print("crib gives", len(mapping), "of 64 alphabet slots:",
          " ".join(f"{k}->{v}" for k, v in sorted(mapping.items())))
    
    # 4. Full alphabet recovered from the binary (the 64-char string near the decoder)
    CUSTOM = b"ZYXWVUTSRQPONMLKJIHGFEDCBAzyxwvutsrqponmlkjihgfedcba9876543210+/"
    assert len(set(CUSTOM)) == 64
    plain = base64.b64decode(blob.translate(bytes.maketrans(CUSTOM, STD)) + b"=" * (-len(blob) % 4))
    print("decoded:", plain.decode())
    text
    length 56 | length % 4 = 0 | distinct chars 36 | outside std alphabet: none
    standard decode: 60% printable -> b'\xc7\xb1\x1b\xc2q\x1b(!\x1d\x015=\x00~(\xbd5 \xc57(:s\x1e'
    crib gives 13 of 64 alphabet slots: 0->9 2->7 F->U G->T P->K V->E X->C Z->A c->x d->w m->n w->d y->b
    decoded: server=update.example.com;port=8080;id=lab

    The blob passes every Base64 shape test, yet standard decoding yields binary junk. The crib recovers only 13 of the 64 slots, but they already show the pattern (Z->A, X->C, V->E): each character range is reversed. In a real sample you would confirm this by finding the 64-character table in the binary. Then paste that table into CyberChef's From Base64 alphabet field and check you get the same output.

  4. Brute-force the rolling key. You have seen a decoder of the roll_sub shape (a seed that grows by a constant), but not its constants. Save as roll_brute.py:

    python
    # roll_brute.py - brute-force a rolling ADD key: out[i] = in[i] + seed + step*i
    import string, sys
    
    data = open(sys.argv[1], "rb").read()
    OK = set((string.ascii_letters + string.digits + "=;.:/_-").encode())
    
    def decode(seed: int, step: int) -> bytes:
        return bytes((b - seed - step * i) & 0xFF for i, b in enumerate(data))
    
    hits = []
    for seed in range(256):
        for step in range(256):
            out = decode(seed, step)
            score = sum(c in OK for c in out) / len(out)
            if score == 1.0:
                hits.append((seed, step, out))
    
    print(f"tried 65536 (seed, step) pairs, {len(hits)} fully printable")
    for seed, step, out in hits[:5]:
        print(f"seed=0x{seed:02x} step={step:<3} {out.decode()}")
    print("with crib 'server=':",
          [(hex(s), st) for s, st, o in hits if o.startswith(b"server=")])
    text
    tried 65536 (seed, step) pairs, 1 fully printable
    seed=0x5c step=3   server=update.example.com;port=8080;id=lab
    with crib 'server=': [('0x5c', 3)]

    The whole search takes about 0.3 seconds. Only one of 65,536 candidates decodes every byte to a config-like character.

  5. Cross-check in CyberChef. Paste the XOR blob's hex, apply From Hex and XOR with key 4b1fa237 (Hex). Then paste the Base64 blob, apply From Base64 and replace the alphabet with the custom table. Both should reproduce the config.

Questions to answer: Why did the Hamming test score length 8 slightly better than 4, and what rule protects you from that? The Base64 blob contains only 36 distinct characters. Would you still call it Base64 if it had been 12 characters long, and what else would you check? How would the XOR recovery change if the config had been compressed before encoding? If the rolling key's step depended on the previous plaintext byte, why would the 65,536-candidate brute force stop working?

Key takeaways

  • Encodings hide configs, strings and traffic from casual inspection. Decoding them produces IOCs, protocol knowledge and detection content.
  • Character set, length, padding and entropy classify most blobs: hex tops out at 4 bits per byte, Base64 at 6, and simple substitutions keep the plaintext's histogram shape.
  • In code, decoders are short loops with a two-operand XOR, ADD/SUB, rotate or table lookup. The index arithmetic reveals the key length, and a changing key register reveals a rolling scheme.
  • Known plaintext (MZ, http, field names) turns XOR data into its key stream. Key-length tests also accept multiples of the true length, so take the smallest.
  • Watch for null-preserving XOR, stripped Base64 padding and compression under the encoding. Re-measure after every layer you peel.