Skip to content

Leçon 11.2 · Analyse automatisée & avancée· 55 min

Emulating Code with Unicorn and Speakeasy

Why analysts emulate instead of run, CPU vs OS-level emulators, and how to call a sample's own decryption routine with stubbed APIs.

Cette leçon n’est disponible qu’en anglais pour le moment.

Objectifs

  • Explain what emulation gives an analyst that a VM or debugger does not: no OS side effects, full control, and the ability to run one function in isolation
  • Distinguish CPU emulators (Unicorn) from OS-level emulators (speakeasy, Qiling) by what each provides: memory, loader, API and system-call models, hooks
  • Set up memory, stack, arguments and a return trap to call a single function inside a mapped PE
  • Stub imported APIs and hook execution so a sample's own string-decryption routine runs to completion in an emulator
  • Recognise the fidelity limits and anti-emulation tricks that make emulation fail, and decide when to fall back to a debugger

In Shellcode Analysis you ran your first emulation: a raw blob mapped into Unicorn, started at offset 0, stopped at its own ret, with a write hook that watched a URL decode itself. That blob was deliberately easy. It had no imports, needed no arguments, and it was the whole program.

Real samples are rarely that polite. The routine you care about sits in the middle of a PE, expects arguments in registers, reads tables from .rdata, and calls HeapAlloc or VirtualAlloc halfway through. This lesson is about emulating that: loading an image, calling one function inside it the way its caller would, and supplying fake answers for the operating system it thinks it is talking to. It also covers the other family of tools, emulators that model Windows itself, and the ways both families break.

Why emulate instead of run

The previous lesson ran the sample for real and injected analysis code into it. A debugger in your isolated VM also runs it for real. An emulator does not: it interprets the instructions inside a model of a CPU, and every effect those instructions have lands in memory that belongs to your script.

Debugger in a VMInstrumentation (DBI)Emulator
Real OS underneathYesYesNo, only what you or the tool model
Side effectsReal: files, registry, networkRealNone outside the emulator's memory
Where execution startsThe entry point (or where you patch it)The entry pointAny address you choose, with any register state
Control per instructionBreakpoints, single-steppingCallbacks, but at a costEvery instruction, block and memory access can be hooked
ArchitectureMust match the VMMust match the hostAny architecture the emulator supports
FidelityPerfect: it is the real machineNear perfectOnly as good as the model

Four properties make emulation worth its fidelity cost:

  • No OS side effects. An emulated CreateFileA or connect does whatever your stub says. Nothing is written, nothing leaves the host, and the host OS never sees a single system call from the sample. That makes emulation safe to run in places where you would never detonate a sample, including automated pipelines.
  • Full control. You set every register, write any memory, stop after any instruction, snapshot the state and roll it back. You can force a branch by changing a flag, or skip a function by returning from it early.
  • Run one function in isolation. You do not need the program's entry point, its CRT start-up, its anti-debugging checks or its C2 server. If you know the address and the interface of one function, you can call it directly, a thousand times, with any inputs.
  • Cross-architecture. Unicorn runs ARM, MIPS or PowerPC code on an x86-64 laptop, which is how IoT and router malware is often analysed (see ELF for Malware Analysts).

The price is that you, or the tool, must provide everything the operating system normally provides. That is where the two families of emulator differ.

CPU emulators and OS-level emulators

CPU emulators: Unicorn

Unicorn takes QEMU's CPU emulation and exposes it as a library. It gives you a CPU and nothing else:

  • Memory mapping. An empty address space. You mem_map page-aligned regions with permissions and mem_write bytes into them. There is no loader: a PE's sections go where you put them.
  • Registers. Read and write any register, including segment bases (UC_X86_REG_GS_BASE and UC_X86_REG_FS_BASE), which matters because Windows code reaches its TEB and PEB through gs/fs.
  • Hooks. Callbacks on every instruction (UC_HOOK_CODE), every basic block, memory reads and writes, accesses to unmapped memory, interrupts, and specific instructions such as syscall or cpuid.
  • Execution control. emu_start(begin, until, timeout, count) and emu_stop() from inside a hook.

What it does not have: no PE or ELF loader, no imports, no heap, no TEB or PEB, no exceptions dispatched to SEH handlers, no threads, no system calls. An imported call jumps to whatever address sits in the IAT slot, which in a file on disk is not a function at all. Everything above the CPU is your job, which is exactly why Unicorn is ideal for calling one function whose dependencies you can list.

OS-level emulators: speakeasy and Qiling

An OS-level emulator wraps a CPU emulator (both of these wrap Unicorn) with a model of an operating system:

  • A loader that maps the PE, applies relocations and resolves imports against fake DLLs.
  • Process structures: TEB, PEB, a loader module list, command line and environment, so code that walks the PEB (as API hashing does) finds what it expects.
  • API or system-call models. Handlers that implement CreateFileA, RegSetValueExA, InternetOpenUrlA or NtAllocateVirtualMemory well enough to return plausible values, while recording every call and its arguments.
  • An event report: the API trace, files "written", registry keys touched, network requests attempted, and strings found in memory.

speakeasy (Mandiant) models Windows user mode and kernel mode, so it can also run drivers. Its API handlers are written in Python, it emulates PEs and raw shellcode, and it produces a JSON report. It works without real Windows DLLs; unknown APIs are either reported as unsupported or, if configured, answered with a generic stub.

Qiling models Windows, Linux, macOS, FreeBSD, UEFI and others at the system-call level. For Windows it expects a "rootfs" containing real DLLs from a Windows installation you are licensed to use, and emulates the system calls underneath them. That is more faithful for code that relies on real DLL behaviour, and more setup.

A third approach is worth knowing: Dumpulator emulates code starting from a Windows minidump. You run the sample in your lab until the interesting state exists (keys derived, globals initialised, heap populated), take a dump, and then call functions inside it in an emulator with all that state already in place. It closes the gap between "run it" and "emulate it" neatly.

ToolProvidesBest for
UnicornCPU, memory, registers, hooksCalling one function; decoder stubs; any architecture
speakeasyUnicorn + Windows loader, PEB/TEB, Python API models, reportWhole-program API traces of PEs, shellcode and drivers without Windows
QilingUnicorn + multi-OS loaders and system-call layerLinux/IoT samples; Windows code that needs real DLL behaviour
DumpulatorUnicorn + a minidump's real memory and TEBFunctions that depend on runtime-initialised state

Let the malware decrypt for you

The core analyst use of a CPU emulator is not running whole programs. It is this: the sample contains a routine that turns encrypted bytes into plaintext, and you want every plaintext it can produce. The encrypted strings, C2 list or configuration are in the file; the key and algorithm are in the code. Instead of reading the algorithm and reimplementing it in Python, you call the sample's own function on the sample's own data.

That trade-off is worth stating plainly:

Reimplement the algorithmEmulate the routine
Needs you to understandEvery step of the algorithmIts interface: arguments, output, dependencies
CorrectnessYour bug, or the author's quirk, can diverge silentlyBit-exact, because it is the author's code
Speed per stringNative Python speedThousands of emulated instructions
Robustness across variantsBreaks if the author changes the algorithmBreaks if the interface or dependencies change
Best whenYou will run it on thousands of samples in a pipelineYou need the answer today, or the algorithm is ugly

Reimplementing is the subject of Scripting String Decryption. Emulation is what you reach for when the algorithm is a custom keystream, heavy XOR mixing, or simply more code than you want to read.

The workflow has four steps.

  1. Find the decryptor. It is usually a function called from many places with an index, a pointer to encrypted bytes, or both, whose return value is then passed to APIs such as CreateMutexA or InternetConnectA. Cross-references and data flow find it quickly: follow the argument of an interesting API back to the call that produced it.
  2. Pin down its interface. How many arguments, in which registers or stack slots under the calling convention? Where does the output go: a return value, a caller-provided buffer, a global? How long can it be?
  3. List its dependencies. Does it read globals, and are those globals already in the file or filled in at runtime? Does it call imports? Does it call other internal functions that call imports? Each dependency must be mapped, stubbed or pre-initialised.
  4. Call it for every input and harvest the output, with an instruction limit so a mistake cannot loop forever.

Step 3 is where most attempts fail. A routine that uses a key derived at start-up from the volume serial number, or a table decompressed by an initialiser, will produce garbage if you call it on a fresh image. The fixes are to call the initialiser first, to write the derived value into memory yourself, or to start from a dump (Dumpulator) where it already exists.

Setting up a call by hand

To call a function you must recreate, in the emulator, the exact state its caller would have created. On Windows x64:

text
             higher addresses
   ┌──────────────────────────────┐
   │ stack args 5, 6, ...         │  rsp+0x28 ...
   ├──────────────────────────────┤
   │ shadow space (4 × 8 bytes)   │  rsp+0x08 .. rsp+0x27  reserved by the caller
   ├──────────────────────────────┤
   │ return address = your trap   │  rsp  (rsp % 16 == 8 on entry)
   └──────────────────────────────┘
   rcx = arg1   rdx = arg2   r8 = arg3   r9 = arg4   (xmm0-3 for floats)
  • Map the image at its preferred base. If you map a PE at its ImageBase, every absolute pointer inside it (such as the pointers in a string table) is already correct and you can skip relocations. Map it elsewhere and you must apply the base relocations yourself, as the loader would.
  • Map a stack and point rsp into it with room above for the shadow space and any stack arguments.
  • Push a return address you control. The function will eventually ret to it. Use an address that nothing else uses and pass it as until to emu_start, so emulation stops the moment the function returns. This is the call/ret contract turned into a stop condition.
  • Respect alignment. On entry, rsp must be 8 modulo 16, as it is after a real call. Compiled code that uses movaps on stack slots faults on a misaligned stack, and the error will look like a bug in your emulation.
  • Put arguments where the convention says. In 32-bit code that means pushing them on the stack in reverse order, and knowing whether the function is cdecl or stdcall (x86-32 calling conventions).

Hooking and stubbing calls

When the function calls an import, execution jumps to the address in its IAT slot. In a file mapped straight from disk that slot still holds the RVA of a hint/name entry, not code, so emulation dies with an unmapped fetch (you will see exactly this in the lab). The standard fix is IAT stubbing:

  1. Map a small region of fake "API" addresses, each byte a ret (0xC3).
  2. Write one fake address into each IAT slot, and remember which import it stands for.
  3. Hook execution on that region. When a fake address is hit, look up the import name, read its arguments from registers or the stack, do the minimum the caller needs (usually: set rax, maybe write an output buffer), and let the ret return to the caller.

This is API hooking with the real API removed altogether. Three rules keep stubs honest:

  • Only implement what the caller consumes. GetProcessHeap can return any non-zero value if the only consumer is HeapAlloc, which is also yours.
  • Clean up the stack correctly on x86. 32-bit Windows APIs are stdcall: the callee pops its own arguments. A stub that just returns leaves the stack unbalanced, and the caller crashes several instructions later. x64 has no such issue.
  • Stop on anything unexpected. An unknown import should print its name and halt, not return 0 silently. A wrong return value sends the sample down an error path and you will spend an hour reading why your output is empty.

The same hook mechanism skips internal functions: hook the first instruction of a function you do not want to run (a Sleep wrapper, an anti-analysis check, a network routine), set rax to the value its caller hopes for, pop the return address into rip, and carry on. Direct system calls, which some samples use to avoid user-mode hooks, are handled the same way with an instruction hook on syscall: read the service number from eax and emulate or refuse it.

Whole-program emulation with speakeasy

When the question is "what does this sample do?" rather than "what does this one function return?", an OS-level emulator is faster than writing stubs. You point speakeasy at a PE or a shellcode blob and read the API trace: every Windows call, its arguments and the emulated return value. Decrypted strings show up as arguments, which is often the quickest way to recover a mutex name or C2 URL without finding the decryptor at all.

Its report also lists strings found in emulated memory, files and registry keys the sample touched, and network activity it attempted. The environment is configurable (--dump-default-config prints every option): user and host names, OS version, environment variables, command line, DNS answers, HTTP responses and which APIs are allowed to exist.

Two options matter constantly. --entry-point starts emulation at an RVA of your choice instead of the PE's entry point, which lets you skip start-up code the emulator cannot handle. --modules-functions-always-exist answers any unmodelled API with a generic stub instead of stopping. Both are useful, and both let the emulation drift from reality in ways the lab shows.

Limits and anti-emulation

Every emulator is a model, and every model has edges. The common ones:

  • Unmodelled or mis-modelled APIs. An OS emulator implements the APIs its authors needed. The first unmodelled call stops emulation, or returns a generic value that sends the sample somewhere it would never go on Windows.
  • Missing process state. Fields of the PEB, TEB, KUSER_SHARED_DATA or loader structures that the model leaves empty; code that reads them sees zeros.
  • Exceptions, threads and callbacks. SEH-based control flow, work spread across threads, APCs and window-message loops are hard to emulate and often the first thing to go wrong.
  • Instruction coverage. Rare or recent instructions (some AVX-512, certain system instructions) may be unsupported or subtly wrong.
  • Coverage of paths. An emulator, like a sandbox, only sees the paths that execute. Branches that depend on a date, a command from C2 or a missing file stay dark. The next lesson and the upcoming symbolic execution lesson attack that problem from different angles.

Malware authors know all of this, and some samples test for it deliberately:

  • Timing. Emulation is slow and its clocks are synthetic. Measuring instruction time with rdtsc, or comparing GetTickCount before and after a Sleep (sleep-acceleration detection), exposes a clock that does not behave like a real one.
  • CPU identity. cpuid leaves that report an unusual vendor, a hypervisor bit or missing feature flags.
  • Environment fingerprints. speakeasy's defaults are recognisable: the default user is speakeasy_user and the host name speakeasy_host. Samples also check for recent files, mouse movement and uptime (user-activity checks).
  • API hammering. Thousands of calls to cheap APIs before the payload, to exhaust an emulator's time or API budget. speakeasy has a dedicated api_hammering setting because of this.
  • Behavioural probes. Calling an API with invalid arguments and checking that the error code is exactly what Windows returns. A generic stub fails that test.

The responses are the same you would use in a debugger, only cheaper: hook the check and force its result, change the emulator's configuration, emulate from a later starting point, or accept that this sample needs a real run in the lab. The upcoming Anti-VM and Sandbox Evasion lesson in Module 9 goes deeper into the checks themselves.

Tip: When an emulation stops, read the last few API calls and the faulting instruction before touching anything. Nine times out of ten the answer is a return value the model got wrong, and the fix is one stub.

Lab: let a sample decrypt its own strings

You will build a benign program that stores five strings encrypted with a custom keystream and decrypts them on demand through a function that calls GetProcessHeap and HeapAlloc, exactly the shape of a real string decryptor. You will then (1) call that function directly in Unicorn with stubbed APIs, and (2) run the whole program in speakeasy. Every output below is real, from Python 3.12, Unicorn 2.1.4, pefile 2024.8.26, speakeasy 2.0.0b8 and MinGW-w64 GCC 15.2 on an Apple Silicon Mac. Nothing in the lab contacts the network or does anything harmful.

  1. Set up. Use one virtual environment for everything:

    bash
    python3.12 -m venv venv
    ./venv/bin/pip install "git+https://github.com/mandiant/speakeasy@v2.0.0b8"

    speakeasy 2.x pulls in Unicorn 2, pefile and Capstone. The older 1.5.x release on PyPI pins Unicorn 1.0.2, which on this Apple Silicon Mac died with a bus error on the first emulated instruction — an emulator can fail before the sample ever gets a chance to.

  2. Play the malware author. Save encrypt_strings.py, a stand-in for a builder that encrypts the strings and emits a C header:

    python
    # encrypt_strings.py - build-time helper: encrypt the lab's strings and
    # emit a C header. This plays the role of the malware author's builder.
    SEED = 0x4C414231                      # "LAB1"
    STRINGS = [
        b"update.example.com",
        b"443",
        b"/api/v1/checkin",
        b"Global\\LabMutex-7f3a",
        b"Mozilla/5.0 (LabAgent)",
    ]
    
    def keystream(idx, n):
        k = (SEED ^ (idx * 0x9E3779B9)) & 0xFFFFFFFF
        for _ in range(n):
            k = (k * 1103515245 + 12345) & 0xFFFFFFFF
            yield (k >> 16) & 0xFF
    
    with open("strings_enc.h", "w") as f:
        for i, s in enumerate(STRINGS):
            enc = bytes(b ^ k for b, k in zip(s, keystream(i, len(s))))
            f.write(f"static const unsigned char s{i}[] = {{"
                    + ",".join(f"0x{b:02x}" for b in enc) + "};\n")
        f.write("static const struct enc_str { const unsigned char *data; "
                "unsigned len; } g_strings[] = {\n")
        for i, s in enumerate(STRINGS):
            f.write(f"    {{ s{i}, {len(s)} }},\n")
        f.write("};\n")

    and the program, lab.c:

    c
    /* lab.c - benign stand-in for a sample with an encrypted string table.
     * decrypt_str() is the kind of routine you find in real malware: it takes
     * an index, allocates a buffer with Windows APIs and decrypts into it. */
    #include <windows.h>
    #include <stdio.h>
    #include "strings_enc.h"
    
    #define SEED 0x4C414231u
    
    __attribute__((noinline))
    char *decrypt_str(unsigned idx)
    {
        const struct enc_str *e = &g_strings[idx];
        char *out = HeapAlloc(GetProcessHeap(), 0, e->len + 1);
        unsigned k = SEED ^ (idx * 0x9E3779B9u);
        for (unsigned i = 0; i < e->len; i++) {
            k = k * 1103515245u + 12345u;
            out[i] = e->data[i] ^ (unsigned char)(k >> 16);
        }
        out[e->len] = 0;
        return out;
    }
    
    int main(void)
    {
        char *mutex = decrypt_str(3);
        HANDLE h = CreateMutexA(NULL, FALSE, mutex);
        if (GetLastError() == ERROR_ALREADY_EXISTS)
            return 1;
        char *host = decrypt_str(0), *port = decrypt_str(1), *path = decrypt_str(2);
        printf("would contact https://%s:%s%s\n", host, port, path);
        CloseHandle(h);
        return 0;
    }

    Build it. Symbols are kept so you can check your work; in a stripped sample you would find decrypt_str through its callers, as step 4 shows.

    bash
    ./venv/bin/python encrypt_strings.py
    x86_64-w64-mingw32-gcc -O2 -o lab.exe lab.c

    Under Wine it prints would contact https://update.example.com:443/api/v1/checkin. Note that the fifth string, the user-agent, is never used by main.

  3. Confirm static analysis is blind. Search the binary for the plaintext:

    bash
    strings -n 6 lab.exe | grep -iE "example|mutex|mozilla|checkin"
    text
    CreateMutexA
    __imp_CreateMutexA
    CreateMutexA

    Only the import name matches. None of the five strings is visible, as with a real sample using encrypted strings.

  4. Read the decryptor's interface. Disassemble the caller first:

    bash
    x86_64-w64-mingw32-objdump -d -M intel --no-show-raw-insn lab.exe \
        --start-address=0x140002b40 --stop-address=0x140002bc0
    text
    140002b4b:	mov    ecx,0x3
    140002b50:	call   1400014c0 <decrypt_str>
    140002b55:	xor    edx,edx
    140002b57:	xor    ecx,ecx
    140002b59:	mov    r8,rax
    140002b5c:	call   QWORD PTR [rip+0x571e]        # 140008280 <__imp_CreateMutexA>
    ...
    140002b77:	xor    ecx,ecx
    140002b79:	call   1400014c0 <decrypt_str>
    140002b7e:	mov    ecx,0x1
    140002b83:	mov    rsi,rax
    140002b86:	call   1400014c0 <decrypt_str>

    The pattern is unmistakable: a small constant in ecx (the first argument), a call to 0x1400014c0, and the returned pointer in rax passed straight to CreateMutexA as its third argument (r8, the name). Without symbols, CreateMutexA's cross-references lead you to the same call. So the interface is char *f(unsigned index). The start of the function tells you its dependencies:

    text
    1400014c7:	lea    rdi,[rip+0x2b52]        # 140004020 <g_strings>
    1400014ce:	mov    eax,ecx
    1400014d0:	mov    rbx,rax
    1400014d3:	shl    rax,0x4
    1400014d7:	add    rdi,rax
    1400014da:	mov    esi,DWORD PTR [rdi+0x8]
    1400014dd:	call   QWORD PTR [rip+0x6dbd]        # 1400082a0 <__imp_GetProcessHeap>
    1400014e3:	xor    edx,edx
    1400014e5:	lea    r8d,[rsi+0x1]
    1400014e9:	mov    rcx,rax
    1400014ec:	call   QWORD PTR [rip+0x6db6]        # 1400082a8 <__imp_HeapAlloc>

    The index is scaled by 16 (shl rax,4) into a table at 0x140004020, and the dword at +8 in each entry is a length. Dumping the table confirms 16-byte records of pointer, length:

    text
     140004020 c0400040 01000000 12000000 00000000  .@.@............
     140004030 b7400040 01000000 03000000 00000000  .@.@............
     140004040 a8400040 01000000 0f000000 00000000  .@.@............
     140004050 90400040 01000000 14000000 00000000  .@.@............
     140004060 70400040 01000000 16000000 00000000  p@.@............

    Five entries, each pointing into the image. Dependencies: a read-only table already in the file, and two imports. No runtime-initialised state, so a fresh image is enough.

  5. Write the harness. Save call_decryptor.py. It maps the PE at its preferred base, replaces every IAT slot with a fake address, stubs the two APIs, and calls decrypt_str once per index with a return trap:

    python
    # call_decryptor.py - map lab.exe into Unicorn and call its own
    # decrypt_str(idx) for every table entry, stubbing the two APIs it needs.
    import sys
    import pefile
    from unicorn import Uc, UcError, UC_ARCH_X86, UC_MODE_64, UC_HOOK_CODE
    from unicorn.x86_const import (UC_X86_REG_RAX, UC_X86_REG_RCX, UC_X86_REG_R8,
                                   UC_X86_REG_RSP, UC_X86_REG_RIP)
    
    DECRYPT_RVA = 0x14C0          # decrypt_str, from the disassembly
    N_STRINGS   = 5               # g_strings has 5 entries
    STUB_BASE   = 0x70000000      # fake "API" addresses live here
    HEAP_BASE   = 0x60000000      # bump allocator for HeapAlloc
    STACK_BASE  = 0x50000000
    STACK_SIZE  = 0x10000
    RETURN_TRAP = STUB_BASE + 0xFF0   # fake return address: reaching it = done
    
    def align(x, a=0x1000):
        return (x + a - 1) & ~(a - 1)
    
    pe = pefile.PE("lab.exe")
    base = pe.OPTIONAL_HEADER.ImageBase
    mu = Uc(UC_ARCH_X86, UC_MODE_64)
    
    # 1. Map the image the way the Windows loader would: headers + sections.
    mu.mem_map(base, align(pe.OPTIONAL_HEADER.SizeOfImage))
    mu.mem_write(base, pe.header)
    for s in pe.sections:
        data = s.get_data()
        mu.mem_write(base + s.VirtualAddress, data[:s.Misc_VirtualSize or len(data)])
    
    # 2. Point every IAT slot at a unique stub address and remember its name.
    mu.mem_map(STUB_BASE, 0x1000)
    mu.mem_write(STUB_BASE, b"\xC3" * 0x1000)     # every stub is a bare `ret`
    stubs = {}
    PATCH_IAT = "--no-stubs" not in sys.argv
    for i, imp in enumerate(e for d in pe.DIRECTORY_ENTRY_IMPORT for e in d.imports):
        addr = STUB_BASE + i * 8
        if PATCH_IAT:
            mu.mem_write(imp.address, addr.to_bytes(8, "little"))  # imp.address = IAT slot VA
        stubs[addr] = (imp.name or b"?").decode()
    
    # 3. Stack and heap.
    mu.mem_map(STACK_BASE, STACK_SIZE)
    mu.mem_map(HEAP_BASE, 0x10000)
    heap_next = HEAP_BASE
    
    def on_stub(uc, address, size, _):
        global heap_next
        if address == RETURN_TRAP:
            uc.emu_stop()
            return
        name = stubs.get(address, f"unknown@{address:#x}")
        if name == "GetProcessHeap":
            uc.reg_write(UC_X86_REG_RAX, 0x1234)            # any fake handle
        elif name == "HeapAlloc":
            size_req = uc.reg_read(UC_X86_REG_R8)           # 3rd arg: dwBytes
            uc.reg_write(UC_X86_REG_RAX, heap_next)
            print(f"    [stub] HeapAlloc({size_req:#x}) -> {heap_next:#x}")
            heap_next += align(size_req, 0x10)
        else:
            print(f"    [stub] unexpected API {name} - stopping")
            uc.emu_stop()
        # the stub's `ret` then returns to the caller
    
    mu.hook_add(UC_HOOK_CODE, on_stub, begin=STUB_BASE, end=STUB_BASE + 0xFFF)
    
    def call(func_va, arg0):
        """Call func(arg0) with the Windows x64 convention."""
        rsp = STACK_BASE + STACK_SIZE - 0x100
        rsp -= 8
        mu.mem_write(rsp, RETURN_TRAP.to_bytes(8, "little"))   # return address
        mu.reg_write(UC_X86_REG_RSP, rsp)
        mu.reg_write(UC_X86_REG_RCX, arg0)                     # 1st argument
        mu.emu_start(func_va, RETURN_TRAP, count=100_000)      # instruction cap
        return mu.reg_read(UC_X86_REG_RAX)
    
    def read_cstr(addr, limit=256):
        raw = bytes(mu.mem_read(addr, limit))
        return raw.split(b"\0", 1)[0].decode("latin1")
    
    for idx in range(N_STRINGS):
        print(f"decrypt_str({idx})")
        try:
            ptr = call(base + DECRYPT_RVA, idx)
        except UcError as e:
            rip = mu.reg_read(UC_X86_REG_RIP)
            print(f"    emulation error {e} at {rip:#x}")
            continue
        print(f"    -> {ptr:#x}  {read_cstr(ptr)!r}")

    The stack top is 16-byte aligned; subtracting 8 for the return address leaves rsp at 8 modulo 16, as after a real call, with 0x100 bytes above the return address for the shadow space and any stack arguments.

  6. See what happens without stubs first. The --no-stubs switch leaves the IAT as it is on disk:

    bash
    ./venv/bin/python call_decryptor.py --no-stubs
    text
    decrypt_str(0)
        emulation error Invalid memory fetch (UC_ERR_FETCH_UNMAPPED) at 0x8486
    decrypt_str(1)
        emulation error Invalid memory fetch (UC_ERR_FETCH_UNMAPPED) at 0x8486

    0x8486 is not an address at all. It is the value stored in the GetProcessHeap IAT slot on disk: the RVA of its hint/name entry, which the Windows loader would have overwritten with the real function address. The call qword ptr [rip+...] jumped to it and Unicorn had nothing mapped there. This is the most common first error in PE emulation.

  7. Run it with stubs:

    bash
    ./venv/bin/python call_decryptor.py
    text
    decrypt_str(0)
        [stub] HeapAlloc(0x13) -> 0x60000000
        -> 0x60000000  'update.example.com'
    decrypt_str(1)
        [stub] HeapAlloc(0x4) -> 0x60000020
        -> 0x60000020  '443'
    decrypt_str(2)
        [stub] HeapAlloc(0x10) -> 0x60000030
        -> 0x60000030  '/api/v1/checkin'
    decrypt_str(3)
        [stub] HeapAlloc(0x15) -> 0x60000040
        -> 0x60000040  'Global\\LabMutex-7f3a'
    decrypt_str(4)
        [stub] HeapAlloc(0x17) -> 0x60000060
        -> 0x60000060  'Mozilla/5.0 (LabAgent)'

    All five strings, including the user-agent that main never decrypts. You did not read a single line of the keystream: the sample did the work. Each HeapAlloc size is the string length plus one, which confirms the length field you read from the table.

  8. Now let speakeasy run the whole program:

    bash
    ./venv/bin/speakeasy -t lab.exe -o report.json
    text
    * exec: module_entry
    0x140001151: 'api-ms-win-crt-stdio-l1-1-0.__acrt_iob_func(0x2)' -> 0x2
    0xfeedf15c: module_entry: Caught error: unsupported_api
    ...
    Unsupported API: api-ms-win-crt-stdio-l1-1-0.setvbuf (ret: 0x140001164)
    * Finished emulating

    GCC 15 links against the Universal CRT, and the CRT start-up calls setvbuf, which this speakeasy version does not model. Emulation ends before main is reached. Answering unmodelled APIs with a stub gets past setvbuf, and then goes wrong in a more interesting way:

    bash
    ./venv/bin/speakeasy -t lab.exe -o report.json --modules-functions-always-exist
    text
    0x140001164: 'api-ms-win-crt-stdio-l1-1-0.setvbuf(0x2, 0x0, 0x4, 0x0)' -> 0x1
    0x140001170: 'api-ms-win-crt-runtime-l1-1-0._crt_atexit(0x140001010)' -> None
    0x1400013fd: 'api-ms-win-crt-runtime-l1-1-0.abort(0x140001010, 0x0, 0x4, 0x0)'
    -> 0x1
    ...
    0x1400028fc: 'api-ms-win-crt-stdio-l1-1-0.__stdio_common_vfprintf(0x24, 0x2,
    "runtime error 10\\n")' -> 0x11
    0x14000295d: 'api-ms-win-crt-runtime-l1-1-0._exit(0xff)' -> None

    A generic stub returned a value the CRT treated as failure, and the program took its error path. This is the danger of "always exist": the trace is still real emulation, but of a path Windows would never take.

  9. Skip the start-up code. main is at RVA 0x2b40. Start there:

    bash
    ./venv/bin/speakeasy -t lab.exe -o report.json --entry-point 0x2b40
    text
    * exec: module_entry
    0x140001602: 'api-ms-win-crt-runtime-l1-1-0._crt_atexit(0x140001480)' -> None
    0x140002b4b: 'api-ms-win-crt-runtime-l1-1-0._crt_atexit(0x140001580)' -> None
    0x1400014e3: 'kernel32.GetProcessHeap()' -> 0x89a0
    0x1400014f2: 'kernel32.HeapAlloc(0x89a0, 0x0, 0x15)' -> 0x89c0
    0x140002b62: 'kernel32.CreateMutexA(0x0, 0x0, "Global\\LabMutex-7f3a")' -> 0x220
    0x140002b6b: 'kernel32.GetLastError()' -> 0x0
    0x1400014e3: 'kernel32.GetProcessHeap()' -> 0x89a0
    0x1400014f2: 'kernel32.HeapAlloc(0x89a0, 0x0, 0x13)' -> 0x89e0
    0x1400014e3: 'kernel32.GetProcessHeap()' -> 0x89a0
    0x1400014f2: 'kernel32.HeapAlloc(0x89a0, 0x0, 0x4)' -> 0x8a00
    0x1400014e3: 'kernel32.GetProcessHeap()' -> 0x89a0
    0x1400014f2: 'kernel32.HeapAlloc(0x89a0, 0x0, 0x10)' -> 0x8a10
    0x14000288d: 'api-ms-win-crt-stdio-l1-1-0.__acrt_iob_func(0x1)' -> 0x1
    0x1400028ab: 'api-ms-win-crt-stdio-l1-1-0.__stdio_common_vfprintf(0x24, 0x1,
    "would contact https://update.example.com:443/api/v1/checkin\\n")' -> 0x3c
    0x140002bba: 'kernel32.CloseHandle(0x220)' -> 0x1
    * Finished emulating

    The mutex name appears as an argument to CreateMutexA and the URL as the formatted output, with no harness code at all. speakeasy's GetLastError returned 0, so the "already running" check passed. The report also collects strings it found in emulated memory:

    bash
    python3 -c "import json; print(json.load(open('report.json'))['strings']['in_memory']['ansi'])"
    text
    ['Global\\LabMutex-7f3a', 'update.example.com', '/api/v1/checkin']

    The user-agent is missing: main never decrypts it, so no emulation of main can find it. Starting at main also has a cost you should know: the CRT never ran, so a main that used argc/argv would receive garbage. On a sample that reads its command line, this shortcut fails with an invalid read in the first function that touches argv.

Questions to answer: Why did mapping the image at its ImageBase let you skip relocations, and which values in g_strings would have been wrong otherwise? In step 6, what would the fetch address have been if the first import called had been HeapAlloc? If decrypt_str had derived its seed from GetVolumeInformationA instead of a constant, what would your stub need to return, and how would you find the right value? Which approach recovered more IOCs here, calling the decryptor or emulating main, and why? What would a 32-bit stdcall version of the HeapAlloc stub need to do that yours does not?

What to report

Add emulation results to the triage report as derived evidence, with enough detail that someone else can reproduce them:

  • The routine and its interface: address or RVA, arguments, output location, and how you identified it (for example, "called before every CreateMutexA; index in ecx").
  • Dependencies you supplied: which APIs were stubbed and with what return values, which memory you pre-initialised, and where emulation started.
  • Recovered plaintext: every string or configuration field, marked as "recovered by emulating the sample's decryptor", including values the sample's normal execution never uses.
  • Tool and version: emulator, version and configuration changes (user name, host name, --entry-point, generic stubs), since each one can change what the sample does.
  • Where emulation stopped and why, if it did not reach the behaviour you wanted. That is a finding too: an unsupported API or an anti-emulation check tells the next analyst where to start.

Key takeaways

  • Emulation trades fidelity for safety and control: no OS side effects, any starting point, any register state, hooks on every instruction, and any architecture.
  • CPU emulators such as Unicorn give you a CPU and memory only; OS-level emulators such as speakeasy and Qiling add a loader, process structures and API or system-call models, and report what the sample did.
  • The core analyst use of a CPU emulator is calling a sample's own decryption routine on its own data: find it, pin down its interface, satisfy its dependencies, then call it for every input.
  • Calling a function by hand means mapping the image (ideally at its preferred base), building a correctly aligned stack with a return trap, placing arguments per the calling convention, and stubbing every import it reaches.
  • Stubs should return what the caller consumes, balance the stack on 32-bit stdcall, and stop loudly on anything unexpected.
  • Emulators fail on unmodelled APIs, missing process state, exceptions and threads, and are probed on purpose through timing, cpuid, environment fingerprints and API hammering. When a fix is not one stub away, go back to a debugger in the lab.