Leçon 11.2 · Analyse automatisée & avancée· 55 min
Emulating Code with Unicorn and Speakeasy
Why analysts emulate instead of run, CPU vs OS-level emulators, and how to call a sample's own decryption routine with stubbed APIs.
Cette leçon n’est disponible qu’en anglais pour le moment.
Objectifs
- Explain what emulation gives an analyst that a VM or debugger does not: no OS side effects, full control, and the ability to run one function in isolation
- Distinguish CPU emulators (Unicorn) from OS-level emulators (speakeasy, Qiling) by what each provides: memory, loader, API and system-call models, hooks
- Set up memory, stack, arguments and a return trap to call a single function inside a mapped PE
- Stub imported APIs and hook execution so a sample's own string-decryption routine runs to completion in an emulator
- Recognise the fidelity limits and anti-emulation tricks that make emulation fail, and decide when to fall back to a debugger
In Shellcode Analysis you ran your first
emulation: a raw blob mapped into Unicorn, started at offset 0, stopped at its
own ret, with a write hook that watched a URL decode itself. That blob was
deliberately easy. It had no imports, needed no arguments, and it was the whole
program.
Real samples are rarely that polite. The routine you care about sits in the
middle of a PE, expects arguments in registers, reads tables from .rdata, and
calls HeapAlloc or VirtualAlloc halfway through. This lesson is about
emulating that: loading an image, calling one function inside it the way its
caller would, and supplying fake answers for the operating system it thinks it
is talking to. It also covers the other family of tools, emulators that model
Windows itself, and the ways both families break.
Why emulate instead of run
The previous lesson ran the sample for real and injected analysis code into it. A debugger in your isolated VM also runs it for real. An emulator does not: it interprets the instructions inside a model of a CPU, and every effect those instructions have lands in memory that belongs to your script.
| Debugger in a VM | Instrumentation (DBI) | Emulator | |
|---|---|---|---|
| Real OS underneath | Yes | Yes | No, only what you or the tool model |
| Side effects | Real: files, registry, network | Real | None outside the emulator's memory |
| Where execution starts | The entry point (or where you patch it) | The entry point | Any address you choose, with any register state |
| Control per instruction | Breakpoints, single-stepping | Callbacks, but at a cost | Every instruction, block and memory access can be hooked |
| Architecture | Must match the VM | Must match the host | Any architecture the emulator supports |
| Fidelity | Perfect: it is the real machine | Near perfect | Only as good as the model |
Four properties make emulation worth its fidelity cost:
- No OS side effects. An emulated
CreateFileAorconnectdoes whatever your stub says. Nothing is written, nothing leaves the host, and the host OS never sees a single system call from the sample. That makes emulation safe to run in places where you would never detonate a sample, including automated pipelines. - Full control. You set every register, write any memory, stop after any instruction, snapshot the state and roll it back. You can force a branch by changing a flag, or skip a function by returning from it early.
- Run one function in isolation. You do not need the program's entry point, its CRT start-up, its anti-debugging checks or its C2 server. If you know the address and the interface of one function, you can call it directly, a thousand times, with any inputs.
- Cross-architecture. Unicorn runs ARM, MIPS or PowerPC code on an x86-64 laptop, which is how IoT and router malware is often analysed (see ELF for Malware Analysts).
The price is that you, or the tool, must provide everything the operating system normally provides. That is where the two families of emulator differ.
CPU emulators and OS-level emulators
CPU emulators: Unicorn
Unicorn takes QEMU's CPU emulation and exposes it as a library. It gives you a CPU and nothing else:
- Memory mapping. An empty address space. You
mem_mappage-aligned regions with permissions andmem_writebytes into them. There is no loader: a PE's sections go where you put them. - Registers. Read and write any register, including segment bases
(
UC_X86_REG_GS_BASEandUC_X86_REG_FS_BASE), which matters because Windows code reaches its TEB and PEB throughgs/fs. - Hooks. Callbacks on every instruction (
UC_HOOK_CODE), every basic block, memory reads and writes, accesses to unmapped memory, interrupts, and specific instructions such assyscallorcpuid. - Execution control.
emu_start(begin, until, timeout, count)andemu_stop()from inside a hook.
What it does not have: no PE or ELF loader, no imports, no heap, no TEB or PEB, no exceptions dispatched to SEH handlers, no threads, no system calls. An imported call jumps to whatever address sits in the IAT slot, which in a file on disk is not a function at all. Everything above the CPU is your job, which is exactly why Unicorn is ideal for calling one function whose dependencies you can list.
OS-level emulators: speakeasy and Qiling
An OS-level emulator wraps a CPU emulator (both of these wrap Unicorn) with a model of an operating system:
- A loader that maps the PE, applies relocations and resolves imports against fake DLLs.
- Process structures: TEB, PEB, a loader module list, command line and environment, so code that walks the PEB (as API hashing does) finds what it expects.
- API or system-call models. Handlers that implement
CreateFileA,RegSetValueExA,InternetOpenUrlAorNtAllocateVirtualMemorywell enough to return plausible values, while recording every call and its arguments. - An event report: the API trace, files "written", registry keys touched, network requests attempted, and strings found in memory.
speakeasy (Mandiant) models Windows user mode and kernel mode, so it can also run drivers. Its API handlers are written in Python, it emulates PEs and raw shellcode, and it produces a JSON report. It works without real Windows DLLs; unknown APIs are either reported as unsupported or, if configured, answered with a generic stub.
Qiling models Windows, Linux, macOS, FreeBSD, UEFI and others at the system-call level. For Windows it expects a "rootfs" containing real DLLs from a Windows installation you are licensed to use, and emulates the system calls underneath them. That is more faithful for code that relies on real DLL behaviour, and more setup.
A third approach is worth knowing: Dumpulator emulates code starting from a Windows minidump. You run the sample in your lab until the interesting state exists (keys derived, globals initialised, heap populated), take a dump, and then call functions inside it in an emulator with all that state already in place. It closes the gap between "run it" and "emulate it" neatly.
| Tool | Provides | Best for |
|---|---|---|
| Unicorn | CPU, memory, registers, hooks | Calling one function; decoder stubs; any architecture |
| speakeasy | Unicorn + Windows loader, PEB/TEB, Python API models, report | Whole-program API traces of PEs, shellcode and drivers without Windows |
| Qiling | Unicorn + multi-OS loaders and system-call layer | Linux/IoT samples; Windows code that needs real DLL behaviour |
| Dumpulator | Unicorn + a minidump's real memory and TEB | Functions that depend on runtime-initialised state |
Let the malware decrypt for you
The core analyst use of a CPU emulator is not running whole programs. It is this: the sample contains a routine that turns encrypted bytes into plaintext, and you want every plaintext it can produce. The encrypted strings, C2 list or configuration are in the file; the key and algorithm are in the code. Instead of reading the algorithm and reimplementing it in Python, you call the sample's own function on the sample's own data.
That trade-off is worth stating plainly:
| Reimplement the algorithm | Emulate the routine | |
|---|---|---|
| Needs you to understand | Every step of the algorithm | Its interface: arguments, output, dependencies |
| Correctness | Your bug, or the author's quirk, can diverge silently | Bit-exact, because it is the author's code |
| Speed per string | Native Python speed | Thousands of emulated instructions |
| Robustness across variants | Breaks if the author changes the algorithm | Breaks if the interface or dependencies change |
| Best when | You will run it on thousands of samples in a pipeline | You need the answer today, or the algorithm is ugly |
Reimplementing is the subject of Scripting String Decryption. Emulation is what you reach for when the algorithm is a custom keystream, heavy XOR mixing, or simply more code than you want to read.
The workflow has four steps.
- Find the decryptor. It is usually a function called from many places
with an index, a pointer to encrypted bytes, or both, whose return value is
then passed to APIs such as
CreateMutexAorInternetConnectA. Cross-references and data flow find it quickly: follow the argument of an interesting API back to the call that produced it. - Pin down its interface. How many arguments, in which registers or stack slots under the calling convention? Where does the output go: a return value, a caller-provided buffer, a global? How long can it be?
- List its dependencies. Does it read globals, and are those globals already in the file or filled in at runtime? Does it call imports? Does it call other internal functions that call imports? Each dependency must be mapped, stubbed or pre-initialised.
- Call it for every input and harvest the output, with an instruction limit so a mistake cannot loop forever.
Step 3 is where most attempts fail. A routine that uses a key derived at start-up from the volume serial number, or a table decompressed by an initialiser, will produce garbage if you call it on a fresh image. The fixes are to call the initialiser first, to write the derived value into memory yourself, or to start from a dump (Dumpulator) where it already exists.
Setting up a call by hand
To call a function you must recreate, in the emulator, the exact state its caller would have created. On Windows x64:
higher addresses
┌──────────────────────────────┐
│ stack args 5, 6, ... │ rsp+0x28 ...
├──────────────────────────────┤
│ shadow space (4 × 8 bytes) │ rsp+0x08 .. rsp+0x27 reserved by the caller
├──────────────────────────────┤
│ return address = your trap │ rsp (rsp % 16 == 8 on entry)
└──────────────────────────────┘
rcx = arg1 rdx = arg2 r8 = arg3 r9 = arg4 (xmm0-3 for floats)- Map the image at its preferred base. If you map a PE at its
ImageBase, every absolute pointer inside it (such as the pointers in a string table) is already correct and you can skip relocations. Map it elsewhere and you must apply the base relocations yourself, as the loader would. - Map a stack and point
rspinto it with room above for the shadow space and any stack arguments. - Push a return address you control. The function will eventually
retto it. Use an address that nothing else uses and pass it asuntiltoemu_start, so emulation stops the moment the function returns. This is the call/ret contract turned into a stop condition. - Respect alignment. On entry,
rspmust be 8 modulo 16, as it is after a realcall. Compiled code that usesmovapson stack slots faults on a misaligned stack, and the error will look like a bug in your emulation. - Put arguments where the convention says. In 32-bit code that means
pushing them on the stack in reverse order, and
knowing whether the function is
cdeclorstdcall(x86-32 calling conventions).
Hooking and stubbing calls
When the function calls an import, execution jumps to the address in its IAT slot. In a file mapped straight from disk that slot still holds the RVA of a hint/name entry, not code, so emulation dies with an unmapped fetch (you will see exactly this in the lab). The standard fix is IAT stubbing:
- Map a small region of fake "API" addresses, each byte a
ret(0xC3). - Write one fake address into each IAT slot, and remember which import it stands for.
- Hook execution on that region. When a fake address is hit, look up the
import name, read its arguments from registers or the stack, do the minimum
the caller needs (usually: set
rax, maybe write an output buffer), and let theretreturn to the caller.
This is API hooking with the real API removed altogether. Three rules keep stubs honest:
- Only implement what the caller consumes.
GetProcessHeapcan return any non-zero value if the only consumer isHeapAlloc, which is also yours. - Clean up the stack correctly on x86. 32-bit Windows APIs are
stdcall: the callee pops its own arguments. A stub that just returns leaves the stack unbalanced, and the caller crashes several instructions later. x64 has no such issue. - Stop on anything unexpected. An unknown import should print its name and halt, not return 0 silently. A wrong return value sends the sample down an error path and you will spend an hour reading why your output is empty.
The same hook mechanism skips internal functions: hook the first instruction
of a function you do not want to run (a Sleep wrapper, an anti-analysis
check, a network routine), set rax to the value its caller hopes for, pop the
return address into rip, and carry on. Direct system calls, which some
samples use to avoid user-mode hooks, are handled the same way with an
instruction hook on syscall: read the service number from eax and emulate
or refuse it.
Whole-program emulation with speakeasy
When the question is "what does this sample do?" rather than "what does this one function return?", an OS-level emulator is faster than writing stubs. You point speakeasy at a PE or a shellcode blob and read the API trace: every Windows call, its arguments and the emulated return value. Decrypted strings show up as arguments, which is often the quickest way to recover a mutex name or C2 URL without finding the decryptor at all.
Its report also lists strings found in emulated memory, files and registry keys
the sample touched, and network activity it attempted. The environment is
configurable (--dump-default-config prints every option): user and host
names, OS version, environment variables, command line, DNS answers, HTTP
responses and which APIs are allowed to exist.
Two options matter constantly. --entry-point starts emulation at an RVA of
your choice instead of the PE's entry point, which lets you skip start-up code
the emulator cannot handle. --modules-functions-always-exist answers any
unmodelled API with a generic stub instead of stopping. Both are useful, and
both let the emulation drift from reality in ways the lab shows.
Limits and anti-emulation
Every emulator is a model, and every model has edges. The common ones:
- Unmodelled or mis-modelled APIs. An OS emulator implements the APIs its authors needed. The first unmodelled call stops emulation, or returns a generic value that sends the sample somewhere it would never go on Windows.
- Missing process state. Fields of the PEB, TEB,
KUSER_SHARED_DATAor loader structures that the model leaves empty; code that reads them sees zeros. - Exceptions, threads and callbacks. SEH-based control flow, work spread across threads, APCs and window-message loops are hard to emulate and often the first thing to go wrong.
- Instruction coverage. Rare or recent instructions (some AVX-512, certain system instructions) may be unsupported or subtly wrong.
- Coverage of paths. An emulator, like a sandbox, only sees the paths that execute. Branches that depend on a date, a command from C2 or a missing file stay dark. The next lesson and the upcoming symbolic execution lesson attack that problem from different angles.
Malware authors know all of this, and some samples test for it deliberately:
- Timing. Emulation is slow and its clocks are synthetic. Measuring
instruction time with
rdtsc, or comparingGetTickCountbefore and after aSleep(sleep-acceleration detection), exposes a clock that does not behave like a real one. - CPU identity.
cpuidleaves that report an unusual vendor, a hypervisor bit or missing feature flags. - Environment fingerprints. speakeasy's defaults are recognisable: the
default user is
speakeasy_userand the host namespeakeasy_host. Samples also check for recent files, mouse movement and uptime (user-activity checks). - API hammering. Thousands of calls to cheap APIs before the payload, to
exhaust an emulator's time or API budget. speakeasy has a dedicated
api_hammeringsetting because of this. - Behavioural probes. Calling an API with invalid arguments and checking that the error code is exactly what Windows returns. A generic stub fails that test.
The responses are the same you would use in a debugger, only cheaper: hook the check and force its result, change the emulator's configuration, emulate from a later starting point, or accept that this sample needs a real run in the lab. The upcoming Anti-VM and Sandbox Evasion lesson in Module 9 goes deeper into the checks themselves.
Tip: When an emulation stops, read the last few API calls and the faulting instruction before touching anything. Nine times out of ten the answer is a return value the model got wrong, and the fix is one stub.
Lab: let a sample decrypt its own strings
You will build a benign program that stores five strings encrypted with a
custom keystream and decrypts them on demand through a function that calls
GetProcessHeap and HeapAlloc, exactly the shape of a real string
decryptor. You will then (1) call that function directly in Unicorn with
stubbed APIs, and (2) run the whole program in speakeasy. Every output below is
real, from Python 3.12, Unicorn 2.1.4, pefile 2024.8.26, speakeasy 2.0.0b8 and
MinGW-w64 GCC 15.2 on an Apple Silicon Mac. Nothing in the lab contacts the
network or does anything harmful.
-
Set up. Use one virtual environment for everything:
bash python3.12 -m venv venv ./venv/bin/pip install "git+https://github.com/mandiant/speakeasy@v2.0.0b8"speakeasy 2.x pulls in Unicorn 2, pefile and Capstone. The older 1.5.x release on PyPI pins Unicorn 1.0.2, which on this Apple Silicon Mac died with a bus error on the first emulated instruction — an emulator can fail before the sample ever gets a chance to.
-
Play the malware author. Save
encrypt_strings.py, a stand-in for a builder that encrypts the strings and emits a C header:python # encrypt_strings.py - build-time helper: encrypt the lab's strings and # emit a C header. This plays the role of the malware author's builder. SEED = 0x4C414231 # "LAB1" STRINGS = [ b"update.example.com", b"443", b"/api/v1/checkin", b"Global\\LabMutex-7f3a", b"Mozilla/5.0 (LabAgent)", ] def keystream(idx, n): k = (SEED ^ (idx * 0x9E3779B9)) & 0xFFFFFFFF for _ in range(n): k = (k * 1103515245 + 12345) & 0xFFFFFFFF yield (k >> 16) & 0xFF with open("strings_enc.h", "w") as f: for i, s in enumerate(STRINGS): enc = bytes(b ^ k for b, k in zip(s, keystream(i, len(s)))) f.write(f"static const unsigned char s{i}[] = {{" + ",".join(f"0x{b:02x}" for b in enc) + "};\n") f.write("static const struct enc_str { const unsigned char *data; " "unsigned len; } g_strings[] = {\n") for i, s in enumerate(STRINGS): f.write(f" {{ s{i}, {len(s)} }},\n") f.write("};\n")and the program,
lab.c:c /* lab.c - benign stand-in for a sample with an encrypted string table. * decrypt_str() is the kind of routine you find in real malware: it takes * an index, allocates a buffer with Windows APIs and decrypts into it. */ #include <windows.h> #include <stdio.h> #include "strings_enc.h" #define SEED 0x4C414231u __attribute__((noinline)) char *decrypt_str(unsigned idx) { const struct enc_str *e = &g_strings[idx]; char *out = HeapAlloc(GetProcessHeap(), 0, e->len + 1); unsigned k = SEED ^ (idx * 0x9E3779B9u); for (unsigned i = 0; i < e->len; i++) { k = k * 1103515245u + 12345u; out[i] = e->data[i] ^ (unsigned char)(k >> 16); } out[e->len] = 0; return out; } int main(void) { char *mutex = decrypt_str(3); HANDLE h = CreateMutexA(NULL, FALSE, mutex); if (GetLastError() == ERROR_ALREADY_EXISTS) return 1; char *host = decrypt_str(0), *port = decrypt_str(1), *path = decrypt_str(2); printf("would contact https://%s:%s%s\n", host, port, path); CloseHandle(h); return 0; }Build it. Symbols are kept so you can check your work; in a stripped sample you would find
decrypt_strthrough its callers, as step 4 shows.bash ./venv/bin/python encrypt_strings.py x86_64-w64-mingw32-gcc -O2 -o lab.exe lab.cUnder Wine it prints
would contact https://update.example.com:443/api/v1/checkin. Note that the fifth string, the user-agent, is never used bymain. -
Confirm static analysis is blind. Search the binary for the plaintext:
bash strings -n 6 lab.exe | grep -iE "example|mutex|mozilla|checkin"text CreateMutexA __imp_CreateMutexA CreateMutexAOnly the import name matches. None of the five strings is visible, as with a real sample using encrypted strings.
-
Read the decryptor's interface. Disassemble the caller first:
bash x86_64-w64-mingw32-objdump -d -M intel --no-show-raw-insn lab.exe \ --start-address=0x140002b40 --stop-address=0x140002bc0text 140002b4b: mov ecx,0x3 140002b50: call 1400014c0 <decrypt_str> 140002b55: xor edx,edx 140002b57: xor ecx,ecx 140002b59: mov r8,rax 140002b5c: call QWORD PTR [rip+0x571e] # 140008280 <__imp_CreateMutexA> ... 140002b77: xor ecx,ecx 140002b79: call 1400014c0 <decrypt_str> 140002b7e: mov ecx,0x1 140002b83: mov rsi,rax 140002b86: call 1400014c0 <decrypt_str>The pattern is unmistakable: a small constant in
ecx(the first argument), a call to0x1400014c0, and the returned pointer inraxpassed straight toCreateMutexAas its third argument (r8, the name). Without symbols,CreateMutexA's cross-references lead you to the same call. So the interface ischar *f(unsigned index). The start of the function tells you its dependencies:text 1400014c7: lea rdi,[rip+0x2b52] # 140004020 <g_strings> 1400014ce: mov eax,ecx 1400014d0: mov rbx,rax 1400014d3: shl rax,0x4 1400014d7: add rdi,rax 1400014da: mov esi,DWORD PTR [rdi+0x8] 1400014dd: call QWORD PTR [rip+0x6dbd] # 1400082a0 <__imp_GetProcessHeap> 1400014e3: xor edx,edx 1400014e5: lea r8d,[rsi+0x1] 1400014e9: mov rcx,rax 1400014ec: call QWORD PTR [rip+0x6db6] # 1400082a8 <__imp_HeapAlloc>The index is scaled by 16 (
shl rax,4) into a table at0x140004020, and the dword at+8in each entry is a length. Dumping the table confirms 16-byte records of pointer, length:text 140004020 c0400040 01000000 12000000 00000000 .@.@............ 140004030 b7400040 01000000 03000000 00000000 .@.@............ 140004040 a8400040 01000000 0f000000 00000000 .@.@............ 140004050 90400040 01000000 14000000 00000000 .@.@............ 140004060 70400040 01000000 16000000 00000000 p@.@............Five entries, each pointing into the image. Dependencies: a read-only table already in the file, and two imports. No runtime-initialised state, so a fresh image is enough.
-
Write the harness. Save
call_decryptor.py. It maps the PE at its preferred base, replaces every IAT slot with a fake address, stubs the two APIs, and callsdecrypt_stronce per index with a return trap:python # call_decryptor.py - map lab.exe into Unicorn and call its own # decrypt_str(idx) for every table entry, stubbing the two APIs it needs. import sys import pefile from unicorn import Uc, UcError, UC_ARCH_X86, UC_MODE_64, UC_HOOK_CODE from unicorn.x86_const import (UC_X86_REG_RAX, UC_X86_REG_RCX, UC_X86_REG_R8, UC_X86_REG_RSP, UC_X86_REG_RIP) DECRYPT_RVA = 0x14C0 # decrypt_str, from the disassembly N_STRINGS = 5 # g_strings has 5 entries STUB_BASE = 0x70000000 # fake "API" addresses live here HEAP_BASE = 0x60000000 # bump allocator for HeapAlloc STACK_BASE = 0x50000000 STACK_SIZE = 0x10000 RETURN_TRAP = STUB_BASE + 0xFF0 # fake return address: reaching it = done def align(x, a=0x1000): return (x + a - 1) & ~(a - 1) pe = pefile.PE("lab.exe") base = pe.OPTIONAL_HEADER.ImageBase mu = Uc(UC_ARCH_X86, UC_MODE_64) # 1. Map the image the way the Windows loader would: headers + sections. mu.mem_map(base, align(pe.OPTIONAL_HEADER.SizeOfImage)) mu.mem_write(base, pe.header) for s in pe.sections: data = s.get_data() mu.mem_write(base + s.VirtualAddress, data[:s.Misc_VirtualSize or len(data)]) # 2. Point every IAT slot at a unique stub address and remember its name. mu.mem_map(STUB_BASE, 0x1000) mu.mem_write(STUB_BASE, b"\xC3" * 0x1000) # every stub is a bare `ret` stubs = {} PATCH_IAT = "--no-stubs" not in sys.argv for i, imp in enumerate(e for d in pe.DIRECTORY_ENTRY_IMPORT for e in d.imports): addr = STUB_BASE + i * 8 if PATCH_IAT: mu.mem_write(imp.address, addr.to_bytes(8, "little")) # imp.address = IAT slot VA stubs[addr] = (imp.name or b"?").decode() # 3. Stack and heap. mu.mem_map(STACK_BASE, STACK_SIZE) mu.mem_map(HEAP_BASE, 0x10000) heap_next = HEAP_BASE def on_stub(uc, address, size, _): global heap_next if address == RETURN_TRAP: uc.emu_stop() return name = stubs.get(address, f"unknown@{address:#x}") if name == "GetProcessHeap": uc.reg_write(UC_X86_REG_RAX, 0x1234) # any fake handle elif name == "HeapAlloc": size_req = uc.reg_read(UC_X86_REG_R8) # 3rd arg: dwBytes uc.reg_write(UC_X86_REG_RAX, heap_next) print(f" [stub] HeapAlloc({size_req:#x}) -> {heap_next:#x}") heap_next += align(size_req, 0x10) else: print(f" [stub] unexpected API {name} - stopping") uc.emu_stop() # the stub's `ret` then returns to the caller mu.hook_add(UC_HOOK_CODE, on_stub, begin=STUB_BASE, end=STUB_BASE + 0xFFF) def call(func_va, arg0): """Call func(arg0) with the Windows x64 convention.""" rsp = STACK_BASE + STACK_SIZE - 0x100 rsp -= 8 mu.mem_write(rsp, RETURN_TRAP.to_bytes(8, "little")) # return address mu.reg_write(UC_X86_REG_RSP, rsp) mu.reg_write(UC_X86_REG_RCX, arg0) # 1st argument mu.emu_start(func_va, RETURN_TRAP, count=100_000) # instruction cap return mu.reg_read(UC_X86_REG_RAX) def read_cstr(addr, limit=256): raw = bytes(mu.mem_read(addr, limit)) return raw.split(b"\0", 1)[0].decode("latin1") for idx in range(N_STRINGS): print(f"decrypt_str({idx})") try: ptr = call(base + DECRYPT_RVA, idx) except UcError as e: rip = mu.reg_read(UC_X86_REG_RIP) print(f" emulation error {e} at {rip:#x}") continue print(f" -> {ptr:#x} {read_cstr(ptr)!r}")The stack top is 16-byte aligned; subtracting 8 for the return address leaves
rspat 8 modulo 16, as after a realcall, with 0x100 bytes above the return address for the shadow space and any stack arguments. -
See what happens without stubs first. The
--no-stubsswitch leaves the IAT as it is on disk:bash ./venv/bin/python call_decryptor.py --no-stubstext decrypt_str(0) emulation error Invalid memory fetch (UC_ERR_FETCH_UNMAPPED) at 0x8486 decrypt_str(1) emulation error Invalid memory fetch (UC_ERR_FETCH_UNMAPPED) at 0x84860x8486is not an address at all. It is the value stored in theGetProcessHeapIAT slot on disk: the RVA of its hint/name entry, which the Windows loader would have overwritten with the real function address. Thecall qword ptr [rip+...]jumped to it and Unicorn had nothing mapped there. This is the most common first error in PE emulation. -
Run it with stubs:
bash ./venv/bin/python call_decryptor.pytext decrypt_str(0) [stub] HeapAlloc(0x13) -> 0x60000000 -> 0x60000000 'update.example.com' decrypt_str(1) [stub] HeapAlloc(0x4) -> 0x60000020 -> 0x60000020 '443' decrypt_str(2) [stub] HeapAlloc(0x10) -> 0x60000030 -> 0x60000030 '/api/v1/checkin' decrypt_str(3) [stub] HeapAlloc(0x15) -> 0x60000040 -> 0x60000040 'Global\\LabMutex-7f3a' decrypt_str(4) [stub] HeapAlloc(0x17) -> 0x60000060 -> 0x60000060 'Mozilla/5.0 (LabAgent)'All five strings, including the user-agent that
mainnever decrypts. You did not read a single line of the keystream: the sample did the work. EachHeapAllocsize is the string length plus one, which confirms the length field you read from the table. -
Now let speakeasy run the whole program:
bash ./venv/bin/speakeasy -t lab.exe -o report.jsontext * exec: module_entry 0x140001151: 'api-ms-win-crt-stdio-l1-1-0.__acrt_iob_func(0x2)' -> 0x2 0xfeedf15c: module_entry: Caught error: unsupported_api ... Unsupported API: api-ms-win-crt-stdio-l1-1-0.setvbuf (ret: 0x140001164) * Finished emulatingGCC 15 links against the Universal CRT, and the CRT start-up calls
setvbuf, which this speakeasy version does not model. Emulation ends beforemainis reached. Answering unmodelled APIs with a stub gets pastsetvbuf, and then goes wrong in a more interesting way:bash ./venv/bin/speakeasy -t lab.exe -o report.json --modules-functions-always-existtext 0x140001164: 'api-ms-win-crt-stdio-l1-1-0.setvbuf(0x2, 0x0, 0x4, 0x0)' -> 0x1 0x140001170: 'api-ms-win-crt-runtime-l1-1-0._crt_atexit(0x140001010)' -> None 0x1400013fd: 'api-ms-win-crt-runtime-l1-1-0.abort(0x140001010, 0x0, 0x4, 0x0)' -> 0x1 ... 0x1400028fc: 'api-ms-win-crt-stdio-l1-1-0.__stdio_common_vfprintf(0x24, 0x2, "runtime error 10\\n")' -> 0x11 0x14000295d: 'api-ms-win-crt-runtime-l1-1-0._exit(0xff)' -> NoneA generic stub returned a value the CRT treated as failure, and the program took its error path. This is the danger of "always exist": the trace is still real emulation, but of a path Windows would never take.
-
Skip the start-up code.
mainis at RVA0x2b40. Start there:bash ./venv/bin/speakeasy -t lab.exe -o report.json --entry-point 0x2b40text * exec: module_entry 0x140001602: 'api-ms-win-crt-runtime-l1-1-0._crt_atexit(0x140001480)' -> None 0x140002b4b: 'api-ms-win-crt-runtime-l1-1-0._crt_atexit(0x140001580)' -> None 0x1400014e3: 'kernel32.GetProcessHeap()' -> 0x89a0 0x1400014f2: 'kernel32.HeapAlloc(0x89a0, 0x0, 0x15)' -> 0x89c0 0x140002b62: 'kernel32.CreateMutexA(0x0, 0x0, "Global\\LabMutex-7f3a")' -> 0x220 0x140002b6b: 'kernel32.GetLastError()' -> 0x0 0x1400014e3: 'kernel32.GetProcessHeap()' -> 0x89a0 0x1400014f2: 'kernel32.HeapAlloc(0x89a0, 0x0, 0x13)' -> 0x89e0 0x1400014e3: 'kernel32.GetProcessHeap()' -> 0x89a0 0x1400014f2: 'kernel32.HeapAlloc(0x89a0, 0x0, 0x4)' -> 0x8a00 0x1400014e3: 'kernel32.GetProcessHeap()' -> 0x89a0 0x1400014f2: 'kernel32.HeapAlloc(0x89a0, 0x0, 0x10)' -> 0x8a10 0x14000288d: 'api-ms-win-crt-stdio-l1-1-0.__acrt_iob_func(0x1)' -> 0x1 0x1400028ab: 'api-ms-win-crt-stdio-l1-1-0.__stdio_common_vfprintf(0x24, 0x1, "would contact https://update.example.com:443/api/v1/checkin\\n")' -> 0x3c 0x140002bba: 'kernel32.CloseHandle(0x220)' -> 0x1 * Finished emulatingThe mutex name appears as an argument to
CreateMutexAand the URL as the formatted output, with no harness code at all. speakeasy'sGetLastErrorreturned 0, so the "already running" check passed. The report also collects strings it found in emulated memory:bash python3 -c "import json; print(json.load(open('report.json'))['strings']['in_memory']['ansi'])"text ['Global\\LabMutex-7f3a', 'update.example.com', '/api/v1/checkin']The user-agent is missing:
mainnever decrypts it, so no emulation ofmaincan find it. Starting atmainalso has a cost you should know: the CRT never ran, so amainthat usedargc/argvwould receive garbage. On a sample that reads its command line, this shortcut fails with an invalid read in the first function that touchesargv.
Questions to answer: Why did mapping the image at its ImageBase let you
skip relocations, and which values in g_strings would have been wrong
otherwise? In step 6, what would the fetch address have been if the first
import called had been HeapAlloc? If decrypt_str had derived its seed from
GetVolumeInformationA instead of a constant, what would your stub need to
return, and how would you find the right value? Which approach recovered more
IOCs here, calling the decryptor or emulating main, and why? What would a
32-bit stdcall version of the HeapAlloc stub need to do that yours does
not?
What to report
Add emulation results to the triage report as derived evidence, with enough detail that someone else can reproduce them:
- The routine and its interface: address or RVA, arguments, output
location, and how you identified it (for example, "called before every
CreateMutexA; index inecx"). - Dependencies you supplied: which APIs were stubbed and with what return values, which memory you pre-initialised, and where emulation started.
- Recovered plaintext: every string or configuration field, marked as "recovered by emulating the sample's decryptor", including values the sample's normal execution never uses.
- Tool and version: emulator, version and configuration changes (user
name, host name,
--entry-point, generic stubs), since each one can change what the sample does. - Where emulation stopped and why, if it did not reach the behaviour you wanted. That is a finding too: an unsupported API or an anti-emulation check tells the next analyst where to start.
Key takeaways
- Emulation trades fidelity for safety and control: no OS side effects, any starting point, any register state, hooks on every instruction, and any architecture.
- CPU emulators such as Unicorn give you a CPU and memory only; OS-level emulators such as speakeasy and Qiling add a loader, process structures and API or system-call models, and report what the sample did.
- The core analyst use of a CPU emulator is calling a sample's own decryption routine on its own data: find it, pin down its interface, satisfy its dependencies, then call it for every input.
- Calling a function by hand means mapping the image (ideally at its preferred base), building a correctly aligned stack with a return trap, placing arguments per the calling convention, and stubbing every import it reaches.
- Stubs should return what the caller consumes, balance the stack on 32-bit
stdcall, and stop loudly on anything unexpected. - Emulators fail on unmodelled APIs, missing process state, exceptions and
threads, and are probed on purpose through timing,
cpuid, environment fingerprints and API hammering. When a fix is not one stub away, go back to a debugger in the lab.