Skip to content

Leçon 11.1 · Analyse automatisée & avancée· 50 min

Dynamic Binary Instrumentation

How DBI engines like Pin, DynamoRIO, Frida and TinyInst inject analysis code into a running sample to log API calls, map coverage and catch unpacking.

Cette leçon n’est disponible qu’en anglais pour le moment.

Objectifs

  • Explain what instrumentation observes that a debugger makes tedious, and contrast static with dynamic instrumentation
  • Describe how a DBI engine re-translates code into a code cache and where its overhead comes from
  • Choose between Intel Pin, DynamoRIO, Frida and TinyInst for a given analysis question
  • Hook library calls and record executed basic blocks with Frida Interceptor and Stalker
  • Recognise how malware detects instrumentation and how written-then-executed tracking exposes unpacked code

A debugger answers questions one stop at a time. You set a breakpoint, the process halts, you read registers, you continue. That is perfect for "what is in rcx when this decryption routine returns?" and miserable for "which of these 4,000 functions ran at all?", "log every call to CreateFileW with its arguments for the whole run", or "tell me the first time execution lands in memory the sample wrote itself". Answering those with breakpoints means thousands of stops, a script gluing them together, and a sample that notices the stops.

Binary instrumentation turns the question around. Instead of stopping the program to look, you insert your own analysis code into it, so the program reports on itself as it runs: a counter at every basic block, a logging callback in front of every API call, a check after every memory write. This lesson covers how that insertion works, what it costs, which engines exist, and the three jobs analysts use it for most: API logging, coverage and unpacking. It opens the automated-analysis module; the lessons after it swap real execution for emulation, then build taint tracking and symbolic execution on top.

What instrumentation gives you

Think of instrumentation as a programmable, always-on debugger that never pauses. Its reach is the same — it sees registers, memory and control flow — but it records instead of stopping.

QuestionWith a debuggerWith instrumentation
Which code actually ran?Breakpoint on every block, or trace-mode single-stepping that crawlsRecord every block once as it is translated; dump the set at exit
What did it pass to each API?Conditional breakpoints plus a script, per functionA hook per export that logs arguments and return values
Where did it write, and did it run what it wrote?Hardware watchpoints, four at a timeInstrument every store and every indirect branch
How often does this loop spin?Hit counts, if the debugger supports themA counter increment compiled into the loop

The trade-off is that you must decide in advance what to record, and you pay a runtime cost for every piece of analysis code you insert. A debugger is better for open-ended exploration; instrumentation is better for systematic questions across an entire execution. In practice analysts use both: instrumentation to find the interesting region, the debugger (the dynamic-analysis module's x64dbg lesson, still to come) to look closely once there.

Static versus dynamic instrumentation

There are two places to insert code: into the file before it runs, or into the process while it runs.

Static binary instrumentation

Static binary instrumentation (SBI) rewrites the executable on disk. Inserting new instructions shifts every address after them, which breaks every jump, call and pointer that referenced the old layout. Since you cannot reliably find all of those references in a binary (the problem Linear vs Recursive Disassembly spent a whole lesson on), SBI tools avoid moving code in place. Two strategies dominate:

  • Breakpoint replacement. Overwrite the first byte of each instrumented instruction with int3. When it executes, the trap handler runs the analysis code, emulates or restores the original instruction, and resumes. It is simple, but each trap is a kernel round trip, so instrumenting every block is painfully slow.
  • Trampolines. Copy each function to a new section, instrument the copy freely, and overwrite the start of the original with a jmp to the copy. Direct calls into the old code bounce to the new one; indirect jumps that land mid-function need extra patching. Dyninst is the best-known tool in this family.

Both depend on a disassembly good enough to find the code. Against a packed sample, SBI is close to useless: the code you care about does not exist in the file until it runs.

Dynamic binary instrumentation

Dynamic binary instrumentation (DBI) does the rewriting at run time, on code as it is about to execute. Because it only ever handles code that is really reached, it never needs a complete disassembly, it sees unpacked and self-modifying code in its final form, and it follows code the sample generates at run time. That is why every tool in this lesson is dynamic, and why DBI rather than SBI matters to malware analysts.

How a DBI engine works

The core idea behind Pin, DynamoRIO and Frida's Stalker is the same: the original code never runs. The engine runs instrumented copies of it instead.

  1. The engine gets control before the program's first instruction (by launching the process itself, or injecting into it).
  2. A dispatcher takes the address about to execute and asks: is there a translated copy of the block starting here in the code cache?
  3. If not, a JIT compiler decodes the original instructions up to the next control transfer, weaves in the analysis callbacks the user requested, rewrites the terminating branch so it returns to the dispatcher, and stores the result in the cache.
  4. The translated block runs natively on the CPU. When it ends, control returns to the dispatcher with the next original address, and the cycle repeats.

Two optimisations keep this fast. Once both ends of a direct branch are translated, the engine patches the cached copy to jump straight to the target's cached copy, so hot loops stop touching the dispatcher at all. Indirect branches (ret, jmp rax, calls through a table) cannot be linked this way; their target is only known at run time, so each one goes through a fast lookup from original address to cached address.

The engine also works hard at transparency. The program must believe it is running its own code at its own addresses: a call must push the original return address, reads of the code bytes must return the original bytes, and exceptions must appear to come from original addresses. Every gap in that illusion is a detection opportunity, as you will see below.

What it costs

CostWhere it comes from
Start-upTranslating code the first time it runs; worst for large, short-lived programs
Steady stateDispatcher lookups on indirect branches, plus whatever analysis code you inserted
Analysis codeUsually dominates: a callback on every instruction or every memory access multiplies run time many times over
MemoryThe code cache and the engine's own state live in the target's address space

An engine with no analysis attached is only moderately slower than native. A tool that calls out on every memory access can be one or two orders of magnitude slower. The practical rule: instrument at the coarsest granularity that answers your question (routine rather than block, block rather than instruction), and do the heavy lifting in the callback only when a cheap filter says it matters.

The engines analysts use

EnginePlatformsYou writeStrengthsTypical analyst use
Intel PinWindows, Linux (x86, x64)C++ "pintools"Mature, fine-grained hooks at instruction, block, trace, routine and image levelAPI tracing with tiny_tracer, custom unpackers, research tools
DynamoRIOWindows, Linux, macOS (partial), x86/x64/ARMC clientsOpen source, fast, rich client library; ships drcov and drltraceCoverage (drcov) loaded into IDA/Binary Ninja via Lighthouse or into Ghidra via Cartographer; library-call logs
FridaWindows, macOS, Linux, iOS, AndroidJavaScript (or C) in an injected agentFastest to iterate; Interceptor for function hooks, Stalker for DBI tracing; attaches to running processesQuick API logging, argument tampering, mobile and macOS samples
TinyInstWindows, macOS, LinuxC++Instruments only the modules you name, so the rest runs at native speedCoverage of one module inside a large process; fuzzing with Jackalope

Some notes on choosing:

  • Frida is two tools in one. Interceptor is not DBI at all: it patches the first instructions of a function with a jump to a trampoline, an inline API hook placed with surgical precision. Stalker is the true DBI engine, re-compiling blocks of a followed thread into a cache. Most quick jobs need only Interceptor; reach for Stalker when the question is about control flow rather than calls.
  • Pin and DynamoRIO launch the program under their control from the first instruction, which is what you want for a Windows packer or loader that does its interesting work before any attach could happen. tiny_tracer, a Pin tool written for malware analysis, logs API calls and module transitions to a text file and includes options to blunt common anti-analysis checks.
  • TinyInst works on a different principle from the others: it runs the process under a debugger interface, marks the target module non-executable, and redirects execution into a rewritten copy when the resulting exceptions fire. Everything outside the named module runs untouched, which keeps overhead low when you only care about one DLL.

Whatever you choose, run it inside your isolated analysis VM. Instrumentation observes a real execution; the sample does everything it would do otherwise.

Analyst uses

API call logging

The most common job. Hook the exports the sample imports (or, for a sample that resolves APIs by hash, every export of the relevant DLLs) and log arguments on entry and results on exit. Compared with a system-wide monitor, you get the exact arguments and return values as the sample's own thread sees them, including decrypted strings passed to InternetConnectW or registry paths passed to RegSetValueExW. The dynamic-analysis module's API-tracing lesson will compare this with kernel-level tracing.

Code coverage: which code actually ran

Record every block executed in the sample's own module, then paint that set onto the disassembly. DynamoRIO's drcov writes a log of module-relative block offsets that the Lighthouse plugin highlights in IDA or Binary Ninja. Coverage answers questions that are otherwise slow:

  • Which functions are worth reading, and which are dead or dormant in this run?
  • Which side of a check was taken? Diff the coverage of two runs (with and without a debugger, a VM artefact, a command-line argument) and the blocks that differ point straight at the decision.
  • Which commands of a C2 handler were exercised by the traffic you replayed?

Coverage is per run. A block missing from the set was not executed this time; it is not proof the code is unreachable.

Written-then-executed: catching unpacked code

A packer writes the real payload into memory and then jumps to it. An instrumentation tool can watch both halves of that sentence:

  1. Instrument every memory write and record the destination addresses (in practice, the pages) the sample writes to.
  2. Instrument every control transfer, or simply note every block the engine translates, and check whether its address lies in a written region.
  3. The first time execution enters written memory, you have a candidate original entry point. Dump the region (and the process image) at that moment.

Because a DBI engine must re-translate any block whose code changed, it already detects writes to code it has cached; an unpacker extends that bookkeeping to all memory. The approach handles packers you have never seen, and multi-layer packers simply trigger it several times. False positives are the usual suspects: JIT compilers, runtime-generated thunks, and legitimate self-modifying code. The evasion-and-unpacking module's manual-unpacking lesson shows the same idea done by hand with page permissions and breakpoints; here it is automated.

How malware detects instrumentation

DBI is transparent by design, not by guarantee. Samples that check for it look for:

  • Time. Translation and callbacks make code slower, so rdtsc or GetTickCount deltas around a loop grow. Heavy instrumentation fails timing checks more often than a debugger does.
  • Foreign modules and threads. Pin, DynamoRIO and Frida all live inside the target. Their DLLs or shared objects, their extra threads (Frida's agent runs its own), and memory regions holding the code cache show up when a sample enumerates modules, threads or executable memory.
  • Patched prologues. Inline hooks such as Frida's Interceptor rewrite the first bytes of the hooked function. A sample that compares the start of NtReadVirtualMemory with the copy on disk, or looks for a jmp at an API entry, sees them. Pure DBI does not change the original bytes, so this check catches hooks rather than tracers.
  • Imperfect emulation of corner cases. Self-modifying code on the same page, unusual exception flows, and reading the return address off the stack with non-standard code have all been used to expose DBI engines at various times.

When a sample behaves differently under instrumentation than on a bare VM, treat it like any other anti-debugging branch: diff coverage between the two runs to find the check, then patch or hook it.

Lab: hook libc and trace basic blocks with Frida

You will build a harmless C program that reads an environment variable and a small config file and takes one of two paths, then use Frida to log the library calls and record which basic blocks ran on each path. The outputs below are real, from Frida 17.19.0 and frida-tools 14.10.4 in a Python 3.14 virtual environment on macOS 26 (Apple silicon, arm64) with System Integrity Protection enabled; the program was built with the system Apple clang 21. Nothing is installed system-wide.

  1. Create a working directory and a virtual environment:

    bash
    mkdir m11a && cd m11a
    python3 -m venv venv
    ./venv/bin/pip install frida-tools
  2. Save the target program as labprog.c. It checks LAB_USER, reads lab.cfg, and calls active_path or quiet_path depending on the file:

    c
    /* labprog.c - a harmless program with one environment check and one file check */
    #include <fcntl.h>
    #include <stdio.h>
    #include <stdlib.h>
    #include <string.h>
    #include <unistd.h>
    
    __attribute__((noinline)) static int config_says_active(const char *path) {
        char buf[32] = {0};
        int fd = open(path, O_RDONLY);
        if (fd < 0)
            return 0;
        read(fd, buf, sizeof buf - 1);
        close(fd);
        return strncmp(buf, "mode=active", 11) == 0;
    }
    
    __attribute__((noinline)) static void quiet_path(void) {
        puts("nothing to do");
    }
    
    __attribute__((noinline)) static void active_path(const char *who) {
        printf("active mode for %s\n", who);
    }
    
    int main(void) {
        const char *who = getenv("LAB_USER");
        if (who == NULL)
            who = "nobody";
        if (config_says_active("lab.cfg"))
            active_path(who);
        else
            quiet_path();
        return 0;
    }
    bash
    clang -O1 -g -o labprog labprog.c
    echo "mode=quiet" > lab.cfg

    clang ad-hoc signs the result. That matters on macOS: with SIP enabled, Frida could spawn this binary we built ourselves without root, but Apple's platform binaries and apps built with the hardened runtime are protected, and instrumenting them requires relaxing SIP on a dedicated test machine. On Windows and Linux lab VMs this restriction does not apply.

  3. Start with frida-trace, which generates a logging handler for each function you name. Pass an absolute path: with a relative one, Frida 17 looked for an application identifier and failed.

    bash
    ./venv/bin/frida-trace -f "$PWD/labprog" -i getenv -i open
    text
    getenv: Auto-generated handler at ".../__handlers__/libsystem_c.dylib/getenv.js"
    open: Auto-generated handler at ".../__handlers__/libsystem_kernel.dylib/open.js"
    Started tracing 2 functions. Web UI available at http://localhost:58207/
    nothing to do
               /* TID 0x103 */
       215 ms  getenv(name="LAB_USER")
       215 ms  open(path="lab.cfg", oflag=0x0, ...)
       215 ms  getenv(name="STDBUF")
       215 ms  getenv(name="STDBUF0")
       ...
       215 ms  getenv(name="_STDBUF_E")
    Process terminated

    Two lessons in one trace. First, the two calls the program makes are there with their arguments, including which file it opened. Second, most of the getenv calls are not from labprog.c at all: they come from the C library setting up stdout buffering on the first puts. A real sample's trace has the same noise from its runtime, which is why good hooks filter by caller or argument.

  4. Save hook.js. It hooks getenv and open with Interceptor, logging return values as well as arguments, and when main is entered it asks Stalker to follow the thread, recording the start of every block that belongs to labprog itself:

    js
    // hook.js - log two libc calls and record which basic blocks of main ran
    const app = Process.mainModule;
    
    function libc(name) {
      return Module.getGlobalExportByName(name);
    }
    
    Interceptor.attach(libc('getenv'), {
      onEnter(args) { this.name = args[0].readUtf8String(); },
      onLeave(ret) {
        if (this.name.startsWith('LAB_'))
          console.log(`getenv("${this.name}") -> ${ret.isNull() ? 'NULL' : '"' + ret.readUtf8String() + '"'}`);
      }
    });
    
    Interceptor.attach(libc('open'), {
      onEnter(args) { this.path = args[0].readUtf8String(); },
      onLeave(ret) { console.log(`open("${this.path}") -> fd ${ret.toInt32()}`); }
    });
    
    const blocks = new Set();
    const mainFn = app.getExportByName('main');
    
    Interceptor.attach(mainFn, {
      onEnter() {
        Stalker.follow(this.threadId, {
          transform(iterator) {
            let insn = iterator.next();
            const start = insn.address;
            if (start.compare(app.base) >= 0 && start.compare(app.base.add(app.size)) < 0)
              blocks.add(start.sub(app.base).toInt32());
            do { iterator.keep(); } while ((insn = iterator.next()) !== null);
          }
        });
      },
      onLeave() {
        Stalker.unfollow(this.threadId);
        // label each block with the function (or stub section) that contains it
        const stubs = app.enumerateSections().find(s => s.name === '__stubs');
        const stubOff = stubs.address.sub(app.base).toInt32();
        const funcs = app.enumerateSymbols()
          .filter(s => s.type === 'section' && s.name !== '')
          .map(s => ({ off: s.address.sub(app.base).toInt32(), name: s.name }))
          .filter(f => f.off > 0 && f.off < stubOff)
          .sort((a, b) => a.off - b.off);
        const byFunc = {};
        for (const off of [...blocks].sort((a, b) => a - b)) {
          let name = off >= stubOff ? '(__stubs: calls into libc)' : '?';
          if (off < stubOff) for (const f of funcs) if (f.off <= off) name = f.name;
          (byFunc[name] = byFunc[name] || []).push('0x' + off.toString(16));
        }
        console.log(`\n${blocks.size} unique basic blocks of ${app.name} executed:`);
        for (const [name, offs] of Object.entries(byFunc))
          console.log(`  ${name.padEnd(28)} ${offs.join(' ')}`);
      }
    });

    The transform callback is the JIT step from earlier, exposed to you: Stalker calls it once per block as it compiles the block into its cache, and iterator.keep() copies each original instruction into the translation. Because translation happens once per block, recording the start address there gives the set of unique blocks executed at almost no cost. Adding a callout per block (iterator.putCallout) would give execution counts, at a higher price.

  5. Save a small driver, run_hook.py. The frida -q -f ... -l hook.js command line spawned the program and printed its output, but the process exited before the agent's messages arrived and the log came back empty; waiting for the session's detached event fixes that:

    python
    # run_hook.py - spawn labprog suspended, load hook.js, resume, wait for exit
    import os, threading, frida
    
    done = threading.Event()
    device = frida.get_local_device()
    pid = device.spawn([os.path.abspath("labprog")], env={**os.environ})
    session = device.attach(pid)
    session.on("detached", lambda reason, crash: (print(f"[detached: {reason}]"), done.set()))
    script = session.create_script(open("hook.js").read())
    script.on("message", lambda msg, data: print(msg))
    script.load()
    device.resume(pid)
    done.wait(20)

    Spawning suspended, loading the script, then resuming is the pattern to use for malware too: the hooks are in place before the sample's first instruction.

  6. Run it once on each path:

    bash
    echo "mode=quiet" > lab.cfg
    ./venv/bin/python run_hook.py
    echo "mode=active" > lab.cfg
    LAB_USER=analyst ./venv/bin/python run_hook.py
    text
    nothing to do
    getenv("LAB_USER") -> NULL
    open("lab.cfg") -> fd 3
    
    18 unique basic blocks of labprog executed:
      main                         0x470 0x478 0x480 0x49c 0x4a0
      config_says_active           0x4b0 0x4e8 0x4ec 0x4fc 0x504 0x540 0x558
      quiet_path                   0x594
      (__stubs: calls into libc)   0x5ac 0x5b8 0x5c4 0x5dc 0x5e8
    [detached: process-terminated]
    text
    active mode for analyst
    getenv("LAB_USER") -> "analyst"
    open("lab.cfg") -> fd 3
    
    20 unique basic blocks of labprog executed:
      main                         0x470 0x478 0x480 0x484 0x498 0x4a0
      config_says_active           0x4b0 0x4e8 0x4ec 0x4fc 0x504 0x540 0x558
      active_path                  0x56c 0x588
      (__stubs: calls into libc)   0x5ac 0x5b8 0x5c4 0x5d0 0x5e8
    [detached: process-terminated]

    The program's own output arrives first because Frida delivers agent messages asynchronously. The Interceptor lines give you the environment value and the file descriptor, which a static read of the binary could only guess at.

  7. Diff the two block sets. config_says_active ran the same seven blocks both times, so the decision is not inside it. In main, the quiet run visited 0x49c while the active run visited 0x484 and 0x498 instead; one stub also changed (0x5dc versus 0x5d0, the puts and printf stubs). Confirm against the disassembly:

    bash
    objdump -d --no-show-raw-insn labprog | awk '/^0000000100000460/,/^00000001000004b0/'
    text
    100000474:     	bl	0x1000005b8 <_read+0x1000005b8>
    100000478:     	mov	x19, x0
    10000047c:     	bl	0x1000004b0 <_config_says_active>
    100000480:     	cbz	w0, 0x10000049c <_main+0x3c>
    100000484:     	adrp	x8, 0x100000000 <_read+0x100000000>
    ...
    100000494:     	bl	0x10000056c <_active_path>
    100000498:     	b	0x1000004a0 <_main+0x40>
    10000049c:     	bl	0x100000594 <_quiet_path>
    1000004a0:     	mov	w0, #0x0                ; =0

    The branch is the cbz w0 at 0x480, testing the return value of config_says_active. (objdump labels the stub calls oddly as _read+...; 0x5b8 is the getenv stub, 0x5d0 is printf and 0x5dc is puts.) Two details are worth noticing. Stalker ends a block at every bl (call), so its blocks are smaller than the ones a disassembler draws in a CFG. And main's first block at 0x460 never appears: Interceptor overwrote those first instructions with its own jump, and the relocated copies ran from Frida's trampoline before Stalker took over. Your observation tools change what they observe.

Questions to answer: Which single instruction decides between the two paths, and which two coverage entries prove it? Why do most of the getenv calls in the frida-trace output not come from labprog.c, and how would you filter them in hook.js? If a sample checked the first bytes of getenv for a branch instruction, which of the two Frida mechanisms used here would it detect, and which would it miss? How would you extend hook.js to flag the first block executed from memory the program had written?

Key takeaways

  • Instrumentation inserts analysis code into the program so it reports on itself; it answers whole-run questions (every API call, every block, every write) that a stop-and-look debugger makes tedious.
  • Static instrumentation rewrites the file with int3 traps or trampolines and needs a full disassembly; dynamic instrumentation translates code as it runs, so it handles packed, generated and self-modifying code.
  • DBI engines run translated copies from a code cache, link direct branches, look up indirect ones, and preserve original addresses for transparency; the analysis code you insert usually dominates the overhead.
  • Pin and DynamoRIO suit launch-time Windows and Linux tracing (tiny_tracer, drcov), Frida suits fast scripted hooking across platforms, and TinyInst suits low-overhead coverage of one module.
  • The core analyst uses are API logging, coverage diffs to find which branch ran, and written-then-executed tracking to catch unpacked code.
  • Malware detects instrumentation through timing, foreign modules and threads, patched API prologues and corner-case behaviour; diff coverage between instrumented and bare runs to find the check.