Skip to content

Lesson 6.4 · Dynamic Analysis· 45 min

Tracing API and System Calls

Record a sample's API and system calls instead of stepping through them, cut the noise, read the trace as behaviour, and know where tracing goes blind.

Objectives

  • Choose between API Monitor, x64dbg logging, Procmon stacks, frida-trace, ETW and strace/ltrace for a given question
  • Scope a trace to the calls that matter and filter out runtime and system noise
  • Reconstruct behaviour from a trace, following handles, return values and the Win32-to-native call chain
  • Explain how direct system calls, unhooking and late attachment defeat tracing, and which vantage points survive

Debugging Malware with x64dbg ended with a logging breakpoint that never stops, just records. Generalise it and you get tracing: instrument a set of functions, let the sample run at near full speed, and read the log. A trace answers "what did it call, with what, in what order, and what came back?" for thousands of calls at once.

It sits between the views you already have. Behavioural Monitoring shows effects on the system but not how the sample got there; debugging shows one moment in full. A trace shows the sample's own requests across the whole run. The usual workflow: monitor, trace to explain what the monitor showed, and debug only what the trace cannot explain.

Where a tracer sits

Every tracer intercepts calls at some layer, and the layer decides what it can see and how easily it is fooled. On Windows, a typical file write goes through three:

text
sample code ──► kernel32 / kernelbase    CreateFileW, WriteFile          (Win32 API)
            ──► ntdll                    NtCreateFile, NtWriteFile        (native API)
            ──► syscall instruction ──► kernel                         (system call)
ToolWhere it interceptsSeesBlind to
API MonitorInline hooks inside the target processWin32 and native calls with decoded arguments, structures and return valuesDirect syscalls; unhooked or hand-resolved code paths
x64dbg logging breakpointsDebugger breakpoints (int3)Any address you choose, with any expression you can formatAnything you did not put a breakpoint on; anti-debug checks see it
frida-traceInline hooks from an injected JavaScript agentExports and arbitrary addresses, with scriptable handlersSame as API Monitor; the agent is visible in the process
Process MonitorKernel callbacks and a file-system minifilterFile, registry, process, image-load and network events, with a stack per eventAnything that is not one of those event types
ETW (SilkETW, logman, EDR sensors)Providers in the kernel and in user-mode componentsProcess, image, network, and many subsystem events, from outside the processWhatever the enabled providers do not emit
strace (Linux)ptrace stops at every system callEvery syscall with decoded argumentsLibrary-level logic that makes no syscall
ltrace (Linux)Breakpoints on PLT entries of dynamically linked binariesLibrary calls such as strcmp and getenvStatic binaries; direct calls that bypass the PLT

The split that matters most is in-process versus out-of-process. API Monitor, Frida and debugger breakpoints live in or attach to the sample's address space: rich arguments, but the sample can see and undo them. Procmon, ETW and strace observe from the kernel: coarser, but the sample cannot remove them from user mode. API hooking in the glossary covers the hooking mechanics; how Frida's Interceptor rewrites a function prologue is covered in depth in Dynamic Binary Instrumentation.

The tools in practice

API Monitor

API Monitor is the quickest readable Win32 trace on Windows. Its definitions for thousands of APIs decode flags, structures and error codes. Use the 64-bit build for 64-bit samples: tick functions in the API Filter pane, start the sample with Monitor New Process, and read the Summary pane, one row per call with caller module, arguments, return value and error. The Parameters and Call Stack panes expand the selected call, and its breakpoints can pause before or after a call so you can edit arguments.

x64dbg: logging breakpoints and trace

Breakpoints with break condition 0 and a log text are already a tracer. For instruction-level work, Trace into/over until condition evaluates a log text at every instruction until a stop condition holds (say, RIP leaving the sample's module); TraceSetLogFile writes it to disk. It is slow, so aim it at one function, such as a decoder.

Process Monitor stacks

Procmon's underused tracing feature is the stack: an event's Stack tab shows the call chain at the moment of the operation, kernel frames marked K and user frames U. With symbols configured (Options > Configure Symbols), frames resolve to names such as KernelBase.dll!CreateFileW, and the first frame inside the sample's module gives you the RVA of the function responsible.

frida-trace

frida-trace generates a small JavaScript handler for each function you name, injects Frida's agent, and prints one line per call. It runs the same way on Windows, Linux and macOS:

text
frida-trace -f <program> -i <export glob> [-i ...] [-x <exclude glob>] [-a module!offset]
frida-trace -p <pid> | -n <name>      attach to a running process instead

-i takes globs (-i "Reg*Value*"), -I/-X include or exclude whole modules, and -a hooks a non-exported function by module and offset. The generated handlers under __handlers__/ are meant to be edited: print a return value, dump a buffer, or filter by caller. The lab below does exactly that.

ETW-based tracing

Event Tracing for Windows is the kernel's own logging system, and what EDR sensors consume. Kernel providers such as Microsoft-Windows-Kernel-Process, -Kernel-File and -Kernel-Network emit events whatever the sample does to its own ntdll. SilkETW (Mandiant) subscribes to kernel or user-mode providers, filters, and writes JSON you can search or match with YARA. On Linux, eBPF tools such as Tetragon play the same role with policy-defined kernel probes. The trade-off is granularity: ETW says a file was created and by whom, not which decrypted buffer an internal function received.

strace and ltrace on Linux

For ELF samples (ELF for Malware Analysts), strace records every system call and ltrace every call through the PLT into shared libraries. They answer different questions: strace shows openat, connect and execve; ltrace shows the strcmp against a hard-coded password that never reaches the kernel.

bash
strace -f -tt -s 256 -y -o strace.txt ./sample      # follow forks, timestamps, long strings, fd paths
strace -f -e trace=%file,%network,%process ./sample  # only the syscall classes you care about
ltrace -f -i ./sample                                # library calls with the caller's address

ltrace only sees calls that go through the PLT, so statically linked binaries (common for Go and for IoT malware) give it nothing to hook; strace still works on them.

What to trace, and how to keep noise down

A full trace of everything is unreadable: the C runtime alone makes hundreds of calls before main, and system DLLs call each other constantly. Scope it:

  1. Start from a question. "Where does it get its C2?" means networking, string, and crypto APIs. "How does it persist?" means registry, file, service and task APIs. Your triage imports list (see Reading Capabilities from Imports) is the menu.
  2. Trace calls made by the sample, not by the system. Filter by caller: keep a call only if its return address is inside the sample's module (or in memory it allocated). Most tools can do this: API Monitor's module filter, an x64dbg condition such as mod.user([rsp]), or an if in a Frida handler.
  3. Pick one layer per question. Tracing both CreateFileW and NtCreateFile doubles every entry. Trace Win32 for readability; add the native layer only when you suspect the sample skips Win32.
  4. Skip start-up. Attach after initialisation or ignore everything before the first call from the sample's own code, if you know start-up is not the interesting part.
  5. Log returns as well as arguments. A trace of calls without results cannot tell a failed check from a successful one.

The lab shows the effect of step 2 on a real trace: about 184,000 lines cut to 27 calls, three of which are the program's own behaviour.

Tip: Hand-resolved APIs (GetProcAddress or API hashing) never appear in the import table, so hooks that patch the IAT miss them. Inline hooks, which patch the function itself, catch them wherever the pointer came from. Know which kind your tool uses.

Reading a trace

A trace is a list of requests. Turning it into behaviour is mostly bookkeeping:

  • Follow handles. A handle returned by one call is an argument to the next. CreateFileW returns 0x1a4; later WriteFile(0x1a4, …) and CloseHandle(0x1a4) belong to that file. The same goes for registry keys, sockets and process handles.
  • Read return values and errors. RegOpenKeyExW returning ERROR_FILE_NOT_FOUND, followed by an early exit, is an environment check. A value the sample reads and then compares is often a configuration switch.
  • Group calls into actions. Open, write, close is "dropped a file"; OpenProcess, VirtualAllocEx, WriteProcessMemory, CreateRemoteThread is injection. Name each group and write the timeline in those terms.
  • Watch the arguments for decoded data. Strings that were XOR encrypted in the file arrive in cleartext at the API that uses them. That is often the cheapest way to recover a C2 address.
  • Mind the timing. Long gaps around Sleep or WaitForSingleObject are delays; a tracer that slows the sample can also trip checks that measure elapsed time, such as sleep acceleration detection.

Native API versus Win32 in a trace

When a trace includes both layers, each Win32 call expands into native calls underneath. A few common pairs:

Win32 (kernel32 / kernelbase)Native (ntdll)
CreateFileWNtCreateFile / NtOpenFile
WriteFile, ReadFileNtWriteFile, NtReadFile
VirtualAlloc(Ex), VirtualProtect(Ex)NtAllocateVirtualMemory, NtProtectVirtualMemory
CreateProcessWNtCreateUserProcess
RegSetValueExWNtSetValueKey
CreateMutexWNtCreateMutant
GetEnvironmentVariableARtlQueryEnvironmentVariable_U (no system call at all)

Two consequences. The ANSI (A) functions convert their strings to UTF-16 and call the wide path, so a trace may show both. And a native call that appears without its Win32 parent, with a return address in the sample's module, means the sample called ntdll directly: a deliberate choice worth noting. The upcoming lesson User Mode, Kernel Mode and the Native API (Module 5) covers the layers in detail.

Where tracing breaks

In-process tracers assume the sample goes through the functions they hooked. Malware that expects to be traced does not:

EvasionHow it worksWhat still sees it
Direct system callsThe sample loads the syscall number into eax and executes syscall itself, never entering ntdllKernel-side sources: ETW, Procmon, EDR kernel callbacks
Indirect system callsSame, but jumps to a syscall instruction inside ntdll so the return address looks legitimateAs above; a debugger with a breakpoint on the kernel transition stub
UnhookingMaps a fresh copy of ntdll from disk or \KnownDlls and copies its clean .text over the hooked one, or restores patched prologues byte by byteKernel-side sources; comparing ntdll's code in memory with the file
Hook detectionReads the first bytes of an API and looks for a jmp, or looks for the tracer's module or threadsOut-of-process tools; or hide the tracer and patch the check
Blinding user-mode ETWPatches ntdll!EtwEventWrite in its own processKernel providers, which do not go through that function
Late or missing attachmentThe interesting work happens before you attach, or in a child process you are not tracingSpawn-mode tracing; following children; system-wide sources

Recognise these in the trace itself: a sample whose trace goes quiet while Procmon still shows it writing files is bypassing your hooks. Direct syscalls also leave a static fingerprint: a mov eax, <number> followed by syscall in the sample's own code, which normal applications never contain. Defeating these is part of the Evasion & Unpacking module; for now, the defensive habit is to pair every in-process trace with one kernel-side source for the same run.

Lab: trace a sample from two sides

Part A traces dbglab.exe from Debugging Malware with x64dbg with Wine's relay channel, which logs every call across a DLL boundary. Part B runs frida-trace on a native macOS build of the same logic; Part C gives the Windows VM equivalents. Output is from Wine 11.0, Frida 17.19.0 with frida-tools 14.10.4 (Python 3.14 venv) and Apple clang 21 on macOS 26 arm64, trimmed where marked. Wine reimplements Windows: the program's calls are the same, but everything below kernel32 is Wine's code, not Microsoft's.

Part A: a Win32 trace of dbglab.exe

  1. Run the program under the relay channel and count what you got:

    bash
    LAB_GO=1 WINEDEBUG=-all,+relay wine dbglab.exe > relay.txt 2>&1
    wineserver -w                  # wait for Wine's helper processes to finish writing
    wc -l < relay.txt
    grep 'GetEnvironmentVariableA' relay.txt | head -1 | cut -d: -f1
    grep -c '^0024:' relay.txt
    grep '^0024:Call' relay.txt | grep -c 'ret=14000'
    text
      183696
    0024
    868
    27

    The file holds every Wine process started for the run, so the total varies between runs. Filtering by the program's thread (0024) leaves 868 lines; keeping only calls whose return address is inside the image (ret=14000…) leaves 27. Most are MinGW start-up code. The tail of the list:

    text
    0024:Call ucrtbase._crt_atexit(140001520) ret=1400013ad
    0024:Call KERNEL32.GetEnvironmentVariableA(140004000 "LAB_GO",0031fdb8,00000008) ret=1400014fc
    0024:Call KERNEL32.OutputDebugStringA(0031fdf0 "http://update.example.com/lab") ret=140002abb
    0024:Call ucrtbase.puts(140004007 "sent to debugger") ret=140002ac7
    0024:Call ucrtbase.exit(00000000) ret=14000140e

    The middle three are the program's behaviour; _crt_atexit and exit are start-up code around main. The return addresses match the call sites found statically in the debugging lab: 0x1400014fc follows the call in the gate function, 0x140002abb the URL call in main. The decoded URL, absent from the file, is in the trace in cleartext.

  2. Look underneath one call. The unfiltered lines for the thread show what each Win32 call did in the layer below (excerpt, lines in between removed where marked):

    text
    0024:Call KERNEL32.GetEnvironmentVariableA(140004000 "LAB_GO",0031fdb8,00000008) ret=1400014fc
    0024:Call ntdll.RtlCreateUnicodeStringFromAsciiz(0031fcc0,140004000 "LAB_GO") ret=6fffffc17954
    0024:Call ntdll.RtlQueryEnvironmentVariable_U(00000000,0031fcc0,0031fcd0) ret=6fffffc17982
    ...
    0024:Ret  KERNEL32.GetEnvironmentVariableA() retval=00000001 ret=1400014fc
    0024:Call KERNEL32.OutputDebugStringA(0031fdf0 "http://update.example.com/lab") ret=140002abb
    0024:Call ntdll.RtlInitUnicodeString(0031fa10,6fffffee74b6 L"DBWinMutex") ret=6fffffc05703
    ...
    0024:Call ntdll.NtCreateMutant(0031fa08,00100000,0031fa20,00000000) ret=6fffffc05733
    ...
    0024:Call ntdll.RtlInitUnicodeString(0031fa00,6fffffee74cc L"DBWIN_BUFFER") ret=6fffffc2ed51
    ...
    0024:Call ntdll.NtOpenSection(0031f9f8,00000002,0031fa10) ret=6fffffc2edb1
    0024:Ret  ntdll.NtOpenSection() retval=c0000034 ret=6fffffc2edb1
    ...
    0024:Ret  KERNEL32.OutputDebugStringA() retval=7ffc0000 ret=140002abb
    0024:Call ucrtbase.puts(140004007 "sent to debugger") ret=140002ac7
    ...
    0024:Call KERNEL32.WriteFile(0000000c,0031eb50,00000012,0031eb4c,00000000) ret=6fffffa0334d
    0024:Call ntdll.NtWriteFile(0000000c,00000000,00000000,00000000,0031ea40,0031eb50,6fff00000012,00000000,00000000) ret=6fffffc68908

    The A function converts its string to UTF-16 first. GetEnvironmentVariableA never leaves user mode, so no syscall-level tool would see it. OutputDebugStringA opens the DBWinMutex mutex and tries the DBWIN_BUFFER section that debug-output viewers create (c0000034 is STATUS_OBJECT_NAME_NOT_FOUND: no viewer was running). And the puts text reaches NtWriteFile only at exit, when the runtime flushes.

Part B: frida-trace on a native build

  1. Save tracelab.c, the same logic for macOS or Linux, plus an optional start-up delay so you can practise attaching:

    c
    /* tracelab.c: harmless program for API tracing — reads a setting, decodes a string, writes a file */
    #include <stdio.h>
    #include <stdlib.h>
    #include <string.h>
    #include <unistd.h>
    
    /* "http://update.example.com/lab" XOR 0x5A */
    static const unsigned char ENC_URL[] = {
        0x32, 0x2e, 0x2e, 0x2a, 0x60, 0x75, 0x75, 0x2f, 0x2a, 0x3e, 0x3b, 0x2e,
        0x3f, 0x74, 0x3f, 0x22, 0x3b, 0x37, 0x2a, 0x36, 0x3f, 0x74, 0x39, 0x35,
        0x37, 0x75, 0x36, 0x3b, 0x38
    };
    
    /* external linkage + noinline: keeps clang from specialising or folding the decoder */
    __attribute__((noinline))
    void decode(char *out, const unsigned char *in, size_t n, unsigned char key) {
        for (size_t i = 0; i < n; i++)
            out[i] = (char)(in[i] ^ key);
        out[n] = '\0';
    }
    
    int main(void) {
        char url[64];
        unsigned delay = getenv("LAB_DELAY") ? (unsigned)atoi(getenv("LAB_DELAY")) : 0;
        if (delay) sleep(delay);                 /* gives you time to attach */
    
        decode(url, ENC_URL, sizeof ENC_URL, 0x5A);
        if (getenv("LAB_GO") == NULL) {
            puts("gate closed");
            return 1;
        }
        FILE *f = fopen("tracelab.out", "w");
        if (f) {
            fprintf(f, "%s\n", url);
            fclose(f);
        }
        puts("wrote tracelab.out");
        return 0;
    }
    bash
    clang -O2 -o tracelab tracelab.c
    python3 -m venv venv && ./venv/bin/pip install frida-tools

    Warning: In a first version, decode was static. Clang then specialised it for its only caller and replaced the loop with constant stores of the plaintext URL, so the string sat in the binary, and a hook on decode printed meaningless arguments (n=259, key=0x3) because the specialised function no longer took them. Internal functions do not have to follow the ABI; check the disassembly before trusting a hook on one.

  2. Trace in spawn mode (-f, with an absolute path), naming the libc calls and the program's own decode:

    bash
    LAB_GO=1 ./venv/bin/frida-trace -f "$PWD/tracelab" -i getenv -i fopen -i puts -i 'tracelab!decode'

    Then edit two generated handlers. __handlers__/libsystem_c.dylib/getenv.js should log the result, not only the name:

    js
    defineHandler({
      onEnter(log, args, state) {
        this.name = args[0].readUtf8String();
      },
      onLeave(log, retval, state) {
        const value = retval.isNull() ? "NULL" : `"${retval.readUtf8String()}"`;
        log(`getenv("${this.name}") => ${value}`);
      }
    });

    And __handlers__/tracelab/decode.js should print the output buffer once the function has filled it:

    js
    defineHandler({
      onEnter(log, args, state) {
        this.out = args[0];                       // decode(out, in, n, key)
        log(`decode(n=${args[2].toInt32()}, key=${ptr(args[3]).and(0xff)})`);
      },
      onLeave(log, retval, state) {
        log(`decode => "${this.out.readUtf8String()}"`);
      }
    });
  3. Run again with both paths. For the gate-open run we also added -i write -i __write_nocancel, with this handler in __handlers__/libsystem_kernel.dylib/__write_nocancel.js, to catch the bytes fclose flushes to the file:

    js
    defineHandler({
      onEnter(log, args, state) {
        const n = args[2].toInt32();
        log(`__write_nocancel(fd=${args[0].toInt32()}, ${JSON.stringify(args[1].readUtf8String(n))}, ${n})`);
      }
    });

    Gate open (start-up lines trimmed):

    text
    Started tracing 6 functions. Web UI available at http://localhost:62711/
               /* TID 0x103 */
         7 ms  getenv("LAB_DELAY") => NULL
         7 ms  decode(n=29, key=0x5a)
         7 ms  decode => "http://update.example.com/lab"
         7 ms  getenv("LAB_GO") => "1"
         7 ms  fopen(path="tracelab.out", mode="w")
         7 ms  getenv("STDBUF") => NULL
         7 ms  getenv("STDBUF0") => NULL
         7 ms  getenv("STDBUF1") => NULL
         7 ms  getenv("STDBUF2") => NULL
         7 ms  getenv("_STDBUF_I") => NULL
         7 ms  getenv("_STDBUF_O") => NULL
         7 ms  getenv("_STDBUF_E") => NULL
         7 ms  __write_nocancel(fd=3, "http://update.example.com/lab\n", 30)
         8 ms  puts(s="wrote tracelab.out")
    Process terminated

    Gate closed (no LAB_GO):

    text
               /* TID 0x103 */
         4 ms  getenv("LAB_DELAY") => NULL
         4 ms  decode(n=29, key=0x5a)
         4 ms  decode => "http://update.example.com/lab"
         4 ms  getenv("LAB_GO") => NULL
         4 ms  puts(s="gate closed")
         4 ms     | getenv("STDBUF") => NULL
         4 ms     | getenv("STDBUF0") => NULL
         4 ms     | getenv("STDBUF1") => NULL
         4 ms     | getenv("STDBUF2") => NULL
         4 ms     | getenv("_STDBUF_I") => NULL
         4 ms     | getenv("_STDBUF_O") => NULL
         4 ms     | getenv("_STDBUF_E") => NULL
    Process terminated

    The seven STDBUF lookups are libc configuring its first stream, not the program. | marks calls made inside another traced function (puts); in the first run they are top level because fprintf, untraced, opened the stream. The gate-closed run still decodes the URL before checking, so hooking the internal function recovers the indicator even when the sample refuses to act. One spawn attempt failed with unexpectedly timed out while initializing suspended process; a retry succeeded.

  4. Attach instead of spawning. Start the program with a delay, then attach by name:

    bash
    LAB_GO=1 LAB_DELAY=8 ./tracelab &
    ./venv/bin/frida-trace -n tracelab -i getenv -i fopen -i puts -i 'tracelab!decode'
    text
    Attaching...
    Started tracing 4 functions. Web UI available at http://localhost:62747/
               /* TID 0x103 */
      6267 ms  decode(n=29, key=0x5a)
      6267 ms  decode => "http://update.example.com/lab"
      6267 ms  getenv("LAB_GO") => "1"
      6267 ms  fopen(path="tracelab.out", mode="w")
      ...
      6268 ms  puts(s="wrote tracelab.out")
    Process terminated

    getenv("LAB_DELAY") is missing: it ran before Frida attached. Attaching worked because tracelab is our own ad-hoc signed binary; a protected system binary refuses even spawn mode:

    text
    Spawning `/bin/echo`...
    Failed to attach: unable to access process with pid 92877 from the current user account

    With System Integrity Protection on, Apple's platform binaries cannot be instrumented; Windows and Linux lab VMs have no such restriction for processes you own.

Part C: the same trace on Windows

  1. frida-trace in the VM. Install frida-tools in a virtual environment in the Windows VM, set LAB_GO=1 in a command prompt, and run frida-trace -f C:\lab\dbglab.exe -i "KERNEL32.DLL!OutputDebugStringA" -i "KERNEL32.DLL!GetEnvironmentVariableA". Qualify the module: many kernel32 exports have a KernelBase twin, and an unqualified -i hooks every module that exports the name, so one call can be logged twice. Edit the generated OutputDebugStringA.js to log args[0].readUtf8String().
  2. API Monitor in the VM. Start API Monitor x64, search the API Filter for OutputDebugStringA and GetEnvironmentVariableA and tick them, then Monitor New Process on dbglab.exe. Run it once with LAB_GO unset and once set, and compare the two call lists and their Call Stack panes with the relay trace from Part A.
  3. Cross-check from the kernel side. In the same run, have Procmon record dbglab.exe. Which of the calls in the API Monitor trace produce a Procmon event, and which never do?

Questions to answer: Why does GetEnvironmentVariableA appear in the relay and Frida traces but could never appear in an strace or ETW kernel trace? In Part A, which filter removed more noise: the thread or the return address, and which one would fail on a sample that injects into another process? If dbglab.exe resolved OutputDebugStringA with API hashing, which of the tools in this lab would still log the call? What would a trace of tracelab look like if the sample issued its write as a raw system call?

Key takeaways

  • Tracing records many calls at full speed; debugging examines one moment in depth. Monitor, then trace, then debug what the trace cannot explain.
  • In-process tracers (API Monitor, Frida, logging breakpoints) give rich arguments but can be seen and removed; kernel-side sources (Procmon, ETW, strace, eBPF) are coarser but survive user-mode tricks.
  • Scope every trace to a question, filter by the caller's module, pick one layer, and log return values.
  • Read a trace by following handles, grouping calls into actions, and reading decoded arguments; the API boundary is where encrypted strings appear in cleartext.
  • Win32 calls expand into native ntdll calls, and some never reach the kernel. A native call made directly from the sample's code is a finding.
  • Direct and indirect syscalls, unhooking, hook detection and late attachment defeat in-process tracing; pair every hook-based trace with a kernel-side record of the same run.