Skip to content

Lesson 8.3 · Encoding, Crypto & Signatures· 55 min

Scripting String Decryption

Recover hundreds of encrypted strings at once by finding the decrypt routine, recovering each call's arguments statically, and scripting the algorithm.

Objectives

  • Locate a string-decryption routine using cross-references, call counts and an argument signature
  • Recover each call site's arguments — immediate lengths and RIP-relative data pointers — by walking backwards from the call
  • Reimplement a decryption algorithm in Python and annotate every call site with the plaintext
  • Choose between reimplementation, emulation and debugging when the algorithm resists a clean port
  • Explain what FLOSS automates, and validate recovered plaintext before trusting it

Open a modern commodity sample in a disassembler and the strings window is almost empty: no URLs, no registry paths, no error messages. Instead there is a single small function called from a few hundred places, each call handing it a pointer into one high-entropy blob. This is XOR string encryption at scale, and it is the most common obfuscation you will meet. The author encrypted every string at build time and decrypts each one just before use, so the plaintext exists only for a moment at runtime.

You could set a breakpoint and read a few, but hundreds of call sites is a scripting problem, not a manual one. This lesson turns one decrypt routine into every plaintext, written back into your database. It builds directly on Cross-References and Data Flow: the pivot table there ended with the observation that a function with dozens of callers, each passing a different pointer into one blob, is very often a string decryptor. Here you act on that observation with code.

The shape of the problem

A string decryptor has a recognisable signature, whatever the algorithm inside:

  • One routine, many callers. The author wrote it once and the compiler calls it everywhere. A caller count in the dozens or hundreds is the loudest signal.
  • A data pointer argument. Each call passes the address of an encrypted buffer, usually as a RIP-relative lea into .rdata or .data, sometimes as an index into a table of blobs.
  • A length or a key. Often a second immediate argument: the length of the buffer, an index, or a per-string key.
  • A small body with a loop. xor, add, subtract, rotate, or a table lookup, over a counter. High-entropy input, printable output.

The workflow follows that shape:

  1. Find the routine. Pivot from the blob or from caller count; confirm by reading the body.
  2. Understand its arguments. Which register or stack slot is the ciphertext, which is the length or key, in what calling convention.
  3. Recover each call's arguments statically. For every caller, walk backwards from the call to the instructions that set those registers, and read off the immediate values and pointer targets.
  4. Reimplement the algorithm in Python and decrypt each buffer.
  5. Write the results back into the disassembler as comments and renames, so the plaintext is visible at every use site forever.

Steps 3 to 5 are the script. Steps 1 and 2 are human work you do once.

Finding and understanding the routine

The fastest anchor is the data. Find the high-entropy blob (the strings and entropy work from Module 3 will have flagged it), take xrefs to it, and you land in the decryptor or in the table it indexes. If there is no single blob — each string sits in its own tiny buffer — pivot on caller count instead: sort functions by number of callers and read the top few small ones. Decryptors, allocators and loggers cluster at the top; the decryptor is the one whose body XORs or adds over a loop.

Reading the body answers step 2. You need three facts before you can script anything: which argument is the ciphertext pointer, which (if any) is the length or key, and exactly what the loop computes. Get the loop wrong by one XOR and every recovered string is garbage, so trace it instruction by instruction and check it against one string you decrypt by hand.

Tip: Rename the routine and its parameters the moment you understand them (decrypt_str, p_cipher, len). Every one of the hundreds of call sites then reads meaningfully in the decompiler even before you script anything, and your script's output has names to attach to.

Recovering arguments by walking backwards

The core of the script is the same use-def reasoning from the previous lesson, done mechanically. At each call site the arguments were set a few instructions earlier: on the Windows x64 ABI, the first four integer arguments arrive in rcx, rdx, r8 and r9. So for a call to decrypt(dst, cipher, len), the cipher pointer is whatever last defined rdx and the length is whatever last defined r8d, reading upward from the call.

Two patterns cover most compiler output:

  • Immediate values — mov r8d, 0x12 gives a length of 18 directly.
  • RIP-relative pointers — lea rdx, [rip+0x164e] gives a pointer whose target is address_of_next_instruction + 0x164e. This is the same arithmetic the disassembler uses to compute the xref, and you compute it yourself with Capstone's operand detail.

Walking backwards has to stop somewhere. Bound the search to a small window (a handful of instructions) and never cross another call, because a call may clobber the volatile argument registers you are tracking. If a value comes from somewhere your simple walk cannot see — a loop counter, a value spilled to the stack, a computed index — record the call site as unresolved rather than guessing. A handful of unresolved sites out of hundreds is a normal, honest result; you finish those by hand or by emulation.

Reimplementing the algorithm

Once you have (cipher_pointer, length) for each call, decryption is a direct port of the loop you read. Read the ciphertext bytes out of the file at the pointer — pefile maps a virtual address to file data for you — and apply the same operations the sample applies. Keep the Python faithful to the assembly, including the width of every operation: a byte XOR that the compiler widened to 32-bit registers is still a byte XOR, and truncating to & 0xFF matters when a running index or key exceeds 255.

Writing results back

A list of plaintexts in your terminal is useful once. Plaintext written into the disassembler is useful for the rest of the analysis. Both Ghidra and IDA expose scripting APIs for exactly this, and the concepts map cleanly:

TaskGhidra (Jython/Java or PyGhidra)IDAPython
Iterate call sitesgetReferencesTo(func.getEntryPoint())idautils.CodeRefsTo(ea, 0)
Read an argument valuegetInstructionBefore, decompiler HighFunctionidc.print_operand, ida_ua.decode_insn
Comment the plaintextsetPreComment / setEOLCommentidc.set_cmt(ea, text, 0)
Rename the buffercreateLabel(addr, name, true)idc.set_name(ea, name)
Read bytes from the imagegetBytes(addr, n)idc.get_bytes(ea, n)

The pattern is identical to the external script below; only the API for reading operands and writing comments changes. Running inside the tool has one big advantage — the disassembler has already found the functions and resolved the call sites, so you skip the disassembly step and get the decompiler's data-flow for free when the arguments are not simple immediates.

When reimplementation is hard

Sometimes you cannot cleanly port the algorithm: it calls into the CRT, depends on a large key schedule, mutates itself, or is just long and fiddly. Two fallbacks recover the same plaintext without a faithful reimplementation.

  • Emulate the routine per call site. Load the sample's bytes into a CPU emulator (Unicorn, or Qiling for a fuller environment), set up the arguments the way each call site does, run the function, and read the output buffer from emulated memory. You reuse the real code instead of rewriting it, which is ideal when the algorithm is complex but self-contained. This is the technique behind the upcoming lesson Emulating Code with Unicorn and Speakeasy, and it is closely related to dynamic binary instrumentation.
  • Log it in a debugger. Set a breakpoint on the routine's return, script the debugger to record the output buffer and the return address, and let the sample decrypt its own strings as it runs. This cuts through anything — network-fetched keys, self-modifying code — but only for the paths that actually execute, and it means running the sample, so it belongs in the lab.

Reimplementation, emulation and debugging trade off the same way throughout this path: static is complete but brittle against complexity, dynamic is robust but only sees what runs. For string decryption, reimplement when the loop is short, emulate when it is complex but pure, and debug when it depends on runtime state.

What FLOSS automates

You have met FLOSS already, in Strings and Obfuscated Strings. It automates exactly the emulation fallback: it ranks candidate decoder functions by their features (loops, XOR with non-zero operands, many callers), emulates each with the arguments its call sites would pass, and diffs memory before and after to capture whatever printable data appeared. On many samples FLOSS recovers the strings with no scripting from you, and it should be your first move.

Scripting the decryptor yourself matters when FLOSS cannot: when the decoder does not match its heuristics, when the key comes from outside the function, when there are thousands of call sites you want individually annotated with the plaintext at each use, or when you need the results inside Ghidra rather than in a report. FLOSS finds the strings; a custom script ties each plaintext to the place it is used and to the argument that selected it, which is what you need to understand behaviour rather than just list indicators.

Validating the output

Never trust decrypted output blindly. A wrong key or an off-by-one in the loop produces confident garbage, and half-right output — a correct algorithm on the wrong length — is more dangerous than an obvious failure.

  • Eyeball printability. Real strings are mostly printable and end where you expect. A run of high-bit bytes means the algorithm or the length is wrong.
  • Cross-check one string dynamically. Decrypt one value in a debugger and confirm your script produces the same bytes. One confirmed string validates the whole batch.
  • Sanity-check the count. If the routine has 300 callers and you recovered 40 strings, the other 260 went somewhere — a table you did not model, an argument pattern your walk missed. Account for the gap.
  • Watch the encoding. Output may be UTF-16LE, or wrapped in another layer (Base64, a second XOR). Decode to bytes first, decide the text encoding second.

Lab: recovering every string from a training sample

You will build a harmless sample whose strings are obscured with an XOR-plus-index scheme and decoded through one routine, then write a Python script with pefile and Capstone that finds every call to that routine, recovers each source pointer and length, decrypts the buffers, and prints address to plaintext. The sample does nothing but print; its only purpose is to be taken apart. You need mingw-w64 and Python 3.

  1. Create a working directory and a virtual environment (tested with Capstone 5.0.9 and pefile 2024.8.26):

    bash
    mkdir -p m8b && cd m8b
    python3 -m venv .venv
    .venv/bin/pip install capstone pefile
  2. Generate the encoded data with gen.py. Each string is stored as plain[i] ^ 0x5A ^ i:

    python
    # gen.py: writes lab_data.h with eight index-XOR-encoded strings.
    STRINGS = ["update.example.com", "/api/v1/beacon", "Mozilla/5.0 (LabLab)",
               "LAB-CAMPAIGN-2026", "SOFTWARE\\LabVendor\\Agent", "lab_mutex_9f13",
               "%TEMP%\\lab_agent.tmp", "config.reload"]
    
    def enc(s):
        return bytes((b ^ 0x5A ^ (i & 0xFF)) & 0xFF for i, b in enumerate(s.encode()))
    
    lines = ["/* generated */", "#define STR_COUNT %d" % len(STRINGS)]
    for i, s in enumerate(STRINGS):
        body = ", ".join("0x%02x" % b for b in enc(s))
        lines.append("static const unsigned char s%d[] = { %s };" % (i, body))
    open("lab_data.h", "w").write("\n".join(lines) + "\n")
    print("wrote lab_data.h")
  3. Write lab_sample.c. One decode routine, called once per string with a pointer and a length that are both immediates — the unrolled form a real decryptor takes:

    c
    /* lab_sample.c - harmless: decodes and prints eight obscured strings. */
    #include <stdio.h>
    #include "lab_data.h"
    
    __attribute__((noinline))
    static void decode(char *dst, const unsigned char *src, int len) {
        for (int i = 0; i < len; i++)
            dst[i] = (char)(src[i] ^ 0x5A ^ (i & 0xFF));
        dst[len] = '\0';
    }
    
    #define SHOW(idx) do { decode(buf, s##idx, (int)sizeof(s##idx)); \
                           printf("string[%d] = %s\n", idx, buf); } while (0)
    
    int main(void) {
        char buf[256];
        SHOW(0); SHOW(1); SHOW(2); SHOW(3);
        SHOW(4); SHOW(5); SHOW(6); SHOW(7);
        return 0;
    }

    Build it and confirm it runs (Wine, or a Windows VM):

    bash
    .venv/bin/python gen.py
    x86_64-w64-mingw32-gcc -O2 -s -o lab_sample.exe lab_sample.c
  4. Look at one call site and the routine before scripting. objdump shows the pattern the script will exploit — a length immediate and a RIP-relative pointer just before each call:

    text
    140002a8d:  lea    rcx,[rsp+0xf0]
    140002a95:  mov    r8d,0x12
    140002a9b:  lea    rdx,[rip+0x164e]        # 0x1400040f0
    140002aa2:  call   0x1400014a0

    And the routine itself, whose loop is the algorithm to port — src[i] ^ i ^ 0x5A:

    text
    1400014c0:  movzx  r9d,BYTE PTR [rdx+rax*1]
    1400014c5:  xor    r9d,eax                 ; ^ index
    1400014c8:  xor    r9d,0x5a                ; ^ key
    1400014cc:  mov    BYTE PTR [rcx+rax*1],r9b
    1400014d0:  add    rax,0x1
    1400014d4:  cmp    r8,rax
    1400014d7:  jne    0x1400014c0

    The binary is stripped, so there are no decode or main labels — the script must infer the routine, which is the realistic case.

  5. Write decrypt_strings.py. It infers the decryptor by argument signature (the call target whose sites are most often preceded by a RIP-relative lea rdx and a mov r8d, imm), because ranking by raw caller count alone picks a C-runtime helper here:

    python
    #!/usr/bin/env python3
    """decrypt_strings.py - recover every string, printing address -> plaintext."""
    import argparse
    from collections import Counter
    import pefile
    from capstone import Cs, CS_ARCH_X86, CS_MODE_64
    from capstone.x86 import X86_OP_MEM, X86_OP_IMM, X86_REG_RIP, X86_REG_RDX, X86_REG_R8D
    
    WINDOW = 8  # instructions before a call to search for its arguments
    
    def decode(buf):
        return bytes((b ^ 0x5A ^ (i & 0xFF)) & 0xFF for i, b in enumerate(buf))
    
    def main():
        ap = argparse.ArgumentParser()
        ap.add_argument("exe")
        ap.add_argument("--func", type=lambda s: int(s, 0), default=None)
        args = ap.parse_args()
    
        pe = pefile.PE(args.exe)
        base = pe.OPTIONAL_HEADER.ImageBase
        md = Cs(CS_ARCH_X86, CS_MODE_64)
        md.detail = True
    
        insns = []
        for sec in pe.sections:
            if sec.Characteristics & 0x20000000:          # IMAGE_SCN_MEM_EXECUTE
                code = sec.get_data()[: sec.Misc_VirtualSize]
                insns.extend(md.disasm(code, base + sec.VirtualAddress))
    
        def recover_args(i):
            """Walk back for a RIP-relative rdx (src) and an r8d immediate (len)."""
            src = length = None
            for j in range(i - 1, max(i - 1 - WINDOW, -1), -1):
                p = insns[j]
                if p.mnemonic == "call":                  # do not cross a call
                    break
                if p.mnemonic == "lea" and p.operands[0].reg == X86_REG_RDX:
                    m = p.operands[1]
                    if m.type == X86_OP_MEM and m.mem.base == X86_REG_RIP:
                        src = p.address + p.size + m.mem.disp
                elif p.mnemonic == "mov" and p.operands[0].reg == X86_REG_R8D \
                        and p.operands[1].type == X86_OP_IMM:
                    length = p.operands[1].imm
                if src is not None and length is not None:
                    break
            return src, length
    
        calls = [i for i, ins in enumerate(insns)
                 if ins.mnemonic == "call" and ins.operands
                 and ins.operands[0].type == X86_OP_IMM]
    
        if args.func is not None:
            target = args.func
        else:
            sig = Counter()
            for i in calls:
                s, n = recover_args(i)
                if s is not None and n is not None:
                    sig[insns[i].operands[0].imm] += 1
            target, n = sig.most_common(1)[0]
            print(f"decode routine inferred at 0x{target:x} ({n} signed call sites)")
    
        for i in calls:
            if insns[i].operands[0].imm != target:
                continue
            src, length = recover_args(i)
            if src is None or length is None:
                print(f"  0x{insns[i].address:x}  <could not recover args>")
                continue
            data = pe.get_data(src - base, length)
            print(f"  0x{insns[i].address:x}  src=0x{src:x} len={length:<3} "
                  f"{decode(data).decode('latin-1')!r}")
    
    if __name__ == "__main__":
        main()
  6. Run it:

    bash
    .venv/bin/python decrypt_strings.py lab_sample.exe
    text
    decode routine inferred at 0x1400014a0 (8 signed call sites)
      0x140002aa2  src=0x1400040f0 len=18  'update.example.com'
      0x140002ad2  src=0x1400040d8 len=14  '/api/v1/beacon'
      0x140002b05  src=0x1400040c0 len=20  'Mozilla/5.0 (LabLab)'
      0x140002b38  src=0x1400040a0 len=17  'LAB-CAMPAIGN-2026'
      0x140002b6b  src=0x140004080 len=24  'SOFTWARE\\LabVendor\\Agent'
      0x140002b9e  src=0x140004068 len=14  'lab_mutex_9f13'
      0x140002bd1  src=0x140004050 len=20  '%TEMP%\\lab_agent.tmp'
      0x140002c04  src=0x140004040 len=13  'config.reload'

    Every string is recovered and tied to the call site that uses it. Your addresses will differ with another compiler version; the shape will not.

  7. Take it into the disassembler. Import lab_sample.exe into Ghidra, run the auto-analysis, and adapt the argument walk to the Ghidra API: iterate getReferencesTo the decode routine, read the lea/mov before each call, and call setEOLComment with the plaintext and createLabel on each source buffer. The listing then shows the decrypted string beside every use — the real payoff, since it makes the surrounding logic readable.

Questions to answer: Why does ranking call targets by caller count alone pick the wrong function here, and what does the argument-signature heuristic add? What would recover_args return if the length were computed in a loop counter instead of a mov r8d, imm, and which of the two fallbacks would you reach for? If the strings were UTF-16LE, where in the script would you change the decoding, and how would you confirm you got it right? How would you adapt the script if the cipher pointers were entries in one big table indexed by a small integer argument rather than distinct lea targets?

Key takeaways

  • A string decryptor is one routine with many callers, a data-pointer argument and a small loop; find it by pivoting from the blob or from caller count, and confirm by reading the body.
  • Recover each call's arguments by walking backwards a few instructions from the call, reading immediates and computing RIP-relative pointer targets, and stop rather than guess when a value is computed.
  • Reimplement the loop faithfully — byte widths and index truncation included — and write the plaintext back as comments and labels with the Ghidra or IDA scripting API so every use site becomes readable.
  • When the algorithm resists a clean port, emulate the routine per call site or log its output in a debugger; FLOSS automates the emulation approach and is the right first move.
  • Validate before trusting: check printability, confirm one string dynamically, reconcile the count of recovered strings against the number of call sites, and handle the text encoding as a separate step.