Lesson 6.4 · Dynamic Analysis· 45 min
Tracing API and System Calls
Record a sample's API and system calls instead of stepping through them, cut the noise, read the trace as behaviour, and know where tracing goes blind.
Objectives
- Choose between API Monitor, x64dbg logging, Procmon stacks, frida-trace, ETW and strace/ltrace for a given question
- Scope a trace to the calls that matter and filter out runtime and system noise
- Reconstruct behaviour from a trace, following handles, return values and the Win32-to-native call chain
- Explain how direct system calls, unhooking and late attachment defeat tracing, and which vantage points survive
Debugging Malware with x64dbg ended with a logging breakpoint that never stops, just records. Generalise it and you get tracing: instrument a set of functions, let the sample run at near full speed, and read the log. A trace answers "what did it call, with what, in what order, and what came back?" for thousands of calls at once.
It sits between the views you already have. Behavioural Monitoring shows effects on the system but not how the sample got there; debugging shows one moment in full. A trace shows the sample's own requests across the whole run. The usual workflow: monitor, trace to explain what the monitor showed, and debug only what the trace cannot explain.
Where a tracer sits
Every tracer intercepts calls at some layer, and the layer decides what it can see and how easily it is fooled. On Windows, a typical file write goes through three:
sample code ──► kernel32 / kernelbase CreateFileW, WriteFile (Win32 API)
──► ntdll NtCreateFile, NtWriteFile (native API)
──► syscall instruction ──► kernel (system call)| Tool | Where it intercepts | Sees | Blind to |
|---|---|---|---|
| API Monitor | Inline hooks inside the target process | Win32 and native calls with decoded arguments, structures and return values | Direct syscalls; unhooked or hand-resolved code paths |
| x64dbg logging breakpoints | Debugger breakpoints (int3) | Any address you choose, with any expression you can format | Anything you did not put a breakpoint on; anti-debug checks see it |
| frida-trace | Inline hooks from an injected JavaScript agent | Exports and arbitrary addresses, with scriptable handlers | Same as API Monitor; the agent is visible in the process |
| Process Monitor | Kernel callbacks and a file-system minifilter | File, registry, process, image-load and network events, with a stack per event | Anything that is not one of those event types |
| ETW (SilkETW, logman, EDR sensors) | Providers in the kernel and in user-mode components | Process, image, network, and many subsystem events, from outside the process | Whatever the enabled providers do not emit |
| strace (Linux) | ptrace stops at every system call | Every syscall with decoded arguments | Library-level logic that makes no syscall |
| ltrace (Linux) | Breakpoints on PLT entries of dynamically linked binaries | Library calls such as strcmp and getenv | Static binaries; direct calls that bypass the PLT |
The split that matters most is in-process versus out-of-process. API
Monitor, Frida and debugger breakpoints live in or attach to the sample's
address space: rich arguments, but the sample can see and undo them.
Procmon, ETW and strace observe from the kernel: coarser, but the sample
cannot remove them from user mode. API hooking in
the glossary covers the hooking mechanics; how Frida's Interceptor rewrites
a function prologue is covered in depth in Dynamic Binary
Instrumentation.
The tools in practice
API Monitor
API Monitor is the quickest readable Win32 trace on Windows. Its definitions for thousands of APIs decode flags, structures and error codes. Use the 64-bit build for 64-bit samples: tick functions in the API Filter pane, start the sample with Monitor New Process, and read the Summary pane, one row per call with caller module, arguments, return value and error. The Parameters and Call Stack panes expand the selected call, and its breakpoints can pause before or after a call so you can edit arguments.
x64dbg: logging breakpoints and trace
Breakpoints with break condition 0 and a log text are already a tracer.
For instruction-level work, Trace into/over until condition evaluates a log
text at every instruction until a stop condition holds (say, RIP leaving
the sample's module); TraceSetLogFile writes it to disk. It is slow, so aim
it at one function, such as a decoder.
Process Monitor stacks
Procmon's underused tracing feature is the stack: an event's Stack
tab shows the call chain at the moment of the operation, kernel frames marked
K and user frames U. With symbols configured (Options > Configure
Symbols), frames resolve to names such as KernelBase.dll!CreateFileW, and
the first frame inside the sample's module gives you the RVA of the function
responsible.
frida-trace
frida-trace generates a small JavaScript handler for each function you name,
injects Frida's agent, and prints one line per call. It runs the same way on
Windows, Linux and macOS:
frida-trace -f <program> -i <export glob> [-i ...] [-x <exclude glob>] [-a module!offset]
frida-trace -p <pid> | -n <name> attach to a running process instead-i takes globs (-i "Reg*Value*"), -I/-X include or exclude whole
modules, and -a hooks a non-exported function by module and offset. The
generated handlers under __handlers__/ are meant to be edited: print a
return value, dump a buffer, or filter by caller. The lab below does exactly
that.
ETW-based tracing
Event Tracing for Windows is the kernel's own logging system, and what EDR
sensors consume. Kernel providers such as Microsoft-Windows-Kernel-Process,
-Kernel-File and -Kernel-Network emit events whatever the sample does to
its own ntdll. SilkETW (Mandiant) subscribes to kernel or user-mode
providers, filters, and writes JSON you can search or match with YARA. On
Linux, eBPF tools such as Tetragon play the same role with policy-defined
kernel probes. The trade-off is granularity: ETW says a file was created and
by whom, not which decrypted buffer an internal function received.
strace and ltrace on Linux
For ELF samples (ELF for Malware Analysts),
strace records every system call and ltrace
every call through the PLT into shared libraries. They answer different
questions: strace shows openat, connect and execve; ltrace shows the
strcmp against a hard-coded password that never reaches the kernel.
strace -f -tt -s 256 -y -o strace.txt ./sample # follow forks, timestamps, long strings, fd paths
strace -f -e trace=%file,%network,%process ./sample # only the syscall classes you care about
ltrace -f -i ./sample # library calls with the caller's addressltrace only sees calls that go through the PLT, so statically linked
binaries (common for Go and for IoT malware) give it nothing to hook; strace
still works on them.
What to trace, and how to keep noise down
A full trace of everything is unreadable: the C runtime alone makes hundreds
of calls before main, and system DLLs call each other constantly. Scope it:
- Start from a question. "Where does it get its C2?" means networking, string, and crypto APIs. "How does it persist?" means registry, file, service and task APIs. Your triage imports list (see Reading Capabilities from Imports) is the menu.
- Trace calls made by the sample, not by the system. Filter by caller:
keep a call only if its return address is inside the sample's module (or in
memory it allocated). Most tools can do this: API Monitor's module filter,
an x64dbg condition such as
mod.user([rsp]), or anifin a Frida handler. - Pick one layer per question. Tracing both
CreateFileWandNtCreateFiledoubles every entry. Trace Win32 for readability; add the native layer only when you suspect the sample skips Win32. - Skip start-up. Attach after initialisation or ignore everything before the first call from the sample's own code, if you know start-up is not the interesting part.
- Log returns as well as arguments. A trace of calls without results cannot tell a failed check from a successful one.
The lab shows the effect of step 2 on a real trace: about 184,000 lines cut to 27 calls, three of which are the program's own behaviour.
Tip: Hand-resolved APIs (
GetProcAddressor API hashing) never appear in the import table, so hooks that patch the IAT miss them. Inline hooks, which patch the function itself, catch them wherever the pointer came from. Know which kind your tool uses.
Reading a trace
A trace is a list of requests. Turning it into behaviour is mostly bookkeeping:
- Follow handles. A handle returned by one call is an argument to the
next.
CreateFileWreturns0x1a4; laterWriteFile(0x1a4, …)andCloseHandle(0x1a4)belong to that file. The same goes for registry keys, sockets and process handles. - Read return values and errors.
RegOpenKeyExWreturningERROR_FILE_NOT_FOUND, followed by an early exit, is an environment check. A value the sample reads and then compares is often a configuration switch. - Group calls into actions. Open, write, close is "dropped a file";
OpenProcess,VirtualAllocEx,WriteProcessMemory,CreateRemoteThreadis injection. Name each group and write the timeline in those terms. - Watch the arguments for decoded data. Strings that were XOR encrypted in the file arrive in cleartext at the API that uses them. That is often the cheapest way to recover a C2 address.
- Mind the timing. Long gaps around
SleeporWaitForSingleObjectare delays; a tracer that slows the sample can also trip checks that measure elapsed time, such as sleep acceleration detection.
Native API versus Win32 in a trace
When a trace includes both layers, each Win32 call expands into native calls underneath. A few common pairs:
| Win32 (kernel32 / kernelbase) | Native (ntdll) |
|---|---|
CreateFileW | NtCreateFile / NtOpenFile |
WriteFile, ReadFile | NtWriteFile, NtReadFile |
VirtualAlloc(Ex), VirtualProtect(Ex) | NtAllocateVirtualMemory, NtProtectVirtualMemory |
CreateProcessW | NtCreateUserProcess |
RegSetValueExW | NtSetValueKey |
CreateMutexW | NtCreateMutant |
GetEnvironmentVariableA | RtlQueryEnvironmentVariable_U (no system call at all) |
Two consequences. The ANSI (A) functions convert their strings to
UTF-16 and call the wide path, so a trace may show both. And a native call
that appears without its Win32 parent, with a return address in the
sample's module, means the sample called ntdll directly: a deliberate choice
worth noting. The upcoming lesson User Mode, Kernel Mode and the Native API
(Module 5) covers the layers in detail.
Where tracing breaks
In-process tracers assume the sample goes through the functions they hooked. Malware that expects to be traced does not:
| Evasion | How it works | What still sees it |
|---|---|---|
| Direct system calls | The sample loads the syscall number into eax and executes syscall itself, never entering ntdll | Kernel-side sources: ETW, Procmon, EDR kernel callbacks |
| Indirect system calls | Same, but jumps to a syscall instruction inside ntdll so the return address looks legitimate | As above; a debugger with a breakpoint on the kernel transition stub |
| Unhooking | Maps a fresh copy of ntdll from disk or \KnownDlls and copies its clean .text over the hooked one, or restores patched prologues byte by byte | Kernel-side sources; comparing ntdll's code in memory with the file |
| Hook detection | Reads the first bytes of an API and looks for a jmp, or looks for the tracer's module or threads | Out-of-process tools; or hide the tracer and patch the check |
| Blinding user-mode ETW | Patches ntdll!EtwEventWrite in its own process | Kernel providers, which do not go through that function |
| Late or missing attachment | The interesting work happens before you attach, or in a child process you are not tracing | Spawn-mode tracing; following children; system-wide sources |
Recognise these in the trace itself: a sample whose trace goes quiet while
Procmon still shows it writing files is bypassing your hooks. Direct syscalls
also leave a static fingerprint: a mov eax, <number> followed by
syscall in the sample's own code, which normal applications never contain.
Defeating these is part of the Evasion & Unpacking module; for now, the
defensive habit is to pair every in-process trace with one kernel-side source
for the same run.
Lab: trace a sample from two sides
Part A traces dbglab.exe from Debugging Malware with
x64dbg with Wine's relay channel, which logs
every call across a DLL boundary. Part B runs frida-trace on a native macOS
build of the same logic; Part C gives the Windows VM equivalents. Output is
from Wine 11.0, Frida 17.19.0 with frida-tools 14.10.4 (Python 3.14 venv) and
Apple clang 21 on macOS 26 arm64, trimmed where marked. Wine reimplements
Windows: the program's calls are the same, but everything below kernel32 is
Wine's code, not Microsoft's.
Part A: a Win32 trace of dbglab.exe
-
Run the program under the relay channel and count what you got:
bash LAB_GO=1 WINEDEBUG=-all,+relay wine dbglab.exe > relay.txt 2>&1 wineserver -w # wait for Wine's helper processes to finish writing wc -l < relay.txt grep 'GetEnvironmentVariableA' relay.txt | head -1 | cut -d: -f1 grep -c '^0024:' relay.txt grep '^0024:Call' relay.txt | grep -c 'ret=14000'text 183696 0024 868 27The file holds every Wine process started for the run, so the total varies between runs. Filtering by the program's thread (
0024) leaves 868 lines; keeping only calls whose return address is inside the image (ret=14000…) leaves 27. Most are MinGW start-up code. The tail of the list:text 0024:Call ucrtbase._crt_atexit(140001520) ret=1400013ad 0024:Call KERNEL32.GetEnvironmentVariableA(140004000 "LAB_GO",0031fdb8,00000008) ret=1400014fc 0024:Call KERNEL32.OutputDebugStringA(0031fdf0 "http://update.example.com/lab") ret=140002abb 0024:Call ucrtbase.puts(140004007 "sent to debugger") ret=140002ac7 0024:Call ucrtbase.exit(00000000) ret=14000140eThe middle three are the program's behaviour;
_crt_atexitandexitare start-up code aroundmain. The return addresses match the call sites found statically in the debugging lab:0x1400014fcfollows the call in the gate function,0x140002abbthe URL call inmain. The decoded URL, absent from the file, is in the trace in cleartext. -
Look underneath one call. The unfiltered lines for the thread show what each Win32 call did in the layer below (excerpt, lines in between removed where marked):
text 0024:Call KERNEL32.GetEnvironmentVariableA(140004000 "LAB_GO",0031fdb8,00000008) ret=1400014fc 0024:Call ntdll.RtlCreateUnicodeStringFromAsciiz(0031fcc0,140004000 "LAB_GO") ret=6fffffc17954 0024:Call ntdll.RtlQueryEnvironmentVariable_U(00000000,0031fcc0,0031fcd0) ret=6fffffc17982 ... 0024:Ret KERNEL32.GetEnvironmentVariableA() retval=00000001 ret=1400014fc 0024:Call KERNEL32.OutputDebugStringA(0031fdf0 "http://update.example.com/lab") ret=140002abb 0024:Call ntdll.RtlInitUnicodeString(0031fa10,6fffffee74b6 L"DBWinMutex") ret=6fffffc05703 ... 0024:Call ntdll.NtCreateMutant(0031fa08,00100000,0031fa20,00000000) ret=6fffffc05733 ... 0024:Call ntdll.RtlInitUnicodeString(0031fa00,6fffffee74cc L"DBWIN_BUFFER") ret=6fffffc2ed51 ... 0024:Call ntdll.NtOpenSection(0031f9f8,00000002,0031fa10) ret=6fffffc2edb1 0024:Ret ntdll.NtOpenSection() retval=c0000034 ret=6fffffc2edb1 ... 0024:Ret KERNEL32.OutputDebugStringA() retval=7ffc0000 ret=140002abb 0024:Call ucrtbase.puts(140004007 "sent to debugger") ret=140002ac7 ... 0024:Call KERNEL32.WriteFile(0000000c,0031eb50,00000012,0031eb4c,00000000) ret=6fffffa0334d 0024:Call ntdll.NtWriteFile(0000000c,00000000,00000000,00000000,0031ea40,0031eb50,6fff00000012,00000000,00000000) ret=6fffffc68908The
Afunction converts its string to UTF-16 first.GetEnvironmentVariableAnever leaves user mode, so no syscall-level tool would see it.OutputDebugStringAopens theDBWinMutexmutex and tries theDBWIN_BUFFERsection that debug-output viewers create (c0000034isSTATUS_OBJECT_NAME_NOT_FOUND: no viewer was running). And theputstext reachesNtWriteFileonly at exit, when the runtime flushes.
Part B: frida-trace on a native build
-
Save
tracelab.c, the same logic for macOS or Linux, plus an optional start-up delay so you can practise attaching:c /* tracelab.c: harmless program for API tracing — reads a setting, decodes a string, writes a file */ #include <stdio.h> #include <stdlib.h> #include <string.h> #include <unistd.h> /* "http://update.example.com/lab" XOR 0x5A */ static const unsigned char ENC_URL[] = { 0x32, 0x2e, 0x2e, 0x2a, 0x60, 0x75, 0x75, 0x2f, 0x2a, 0x3e, 0x3b, 0x2e, 0x3f, 0x74, 0x3f, 0x22, 0x3b, 0x37, 0x2a, 0x36, 0x3f, 0x74, 0x39, 0x35, 0x37, 0x75, 0x36, 0x3b, 0x38 }; /* external linkage + noinline: keeps clang from specialising or folding the decoder */ __attribute__((noinline)) void decode(char *out, const unsigned char *in, size_t n, unsigned char key) { for (size_t i = 0; i < n; i++) out[i] = (char)(in[i] ^ key); out[n] = '\0'; } int main(void) { char url[64]; unsigned delay = getenv("LAB_DELAY") ? (unsigned)atoi(getenv("LAB_DELAY")) : 0; if (delay) sleep(delay); /* gives you time to attach */ decode(url, ENC_URL, sizeof ENC_URL, 0x5A); if (getenv("LAB_GO") == NULL) { puts("gate closed"); return 1; } FILE *f = fopen("tracelab.out", "w"); if (f) { fprintf(f, "%s\n", url); fclose(f); } puts("wrote tracelab.out"); return 0; }bash clang -O2 -o tracelab tracelab.c python3 -m venv venv && ./venv/bin/pip install frida-toolsWarning: In a first version,
decodewasstatic. Clang then specialised it for its only caller and replaced the loop with constant stores of the plaintext URL, so the string sat in the binary, and a hook ondecodeprinted meaningless arguments (n=259, key=0x3) because the specialised function no longer took them. Internal functions do not have to follow the ABI; check the disassembly before trusting a hook on one. -
Trace in spawn mode (
-f, with an absolute path), naming the libc calls and the program's owndecode:bash LAB_GO=1 ./venv/bin/frida-trace -f "$PWD/tracelab" -i getenv -i fopen -i puts -i 'tracelab!decode'Then edit two generated handlers.
__handlers__/libsystem_c.dylib/getenv.jsshould log the result, not only the name:js defineHandler({ onEnter(log, args, state) { this.name = args[0].readUtf8String(); }, onLeave(log, retval, state) { const value = retval.isNull() ? "NULL" : `"${retval.readUtf8String()}"`; log(`getenv("${this.name}") => ${value}`); } });And
__handlers__/tracelab/decode.jsshould print the output buffer once the function has filled it:js defineHandler({ onEnter(log, args, state) { this.out = args[0]; // decode(out, in, n, key) log(`decode(n=${args[2].toInt32()}, key=${ptr(args[3]).and(0xff)})`); }, onLeave(log, retval, state) { log(`decode => "${this.out.readUtf8String()}"`); } }); -
Run again with both paths. For the gate-open run we also added
-i write -i __write_nocancel, with this handler in__handlers__/libsystem_kernel.dylib/__write_nocancel.js, to catch the bytesfcloseflushes to the file:js defineHandler({ onEnter(log, args, state) { const n = args[2].toInt32(); log(`__write_nocancel(fd=${args[0].toInt32()}, ${JSON.stringify(args[1].readUtf8String(n))}, ${n})`); } });Gate open (start-up lines trimmed):
text Started tracing 6 functions. Web UI available at http://localhost:62711/ /* TID 0x103 */ 7 ms getenv("LAB_DELAY") => NULL 7 ms decode(n=29, key=0x5a) 7 ms decode => "http://update.example.com/lab" 7 ms getenv("LAB_GO") => "1" 7 ms fopen(path="tracelab.out", mode="w") 7 ms getenv("STDBUF") => NULL 7 ms getenv("STDBUF0") => NULL 7 ms getenv("STDBUF1") => NULL 7 ms getenv("STDBUF2") => NULL 7 ms getenv("_STDBUF_I") => NULL 7 ms getenv("_STDBUF_O") => NULL 7 ms getenv("_STDBUF_E") => NULL 7 ms __write_nocancel(fd=3, "http://update.example.com/lab\n", 30) 8 ms puts(s="wrote tracelab.out") Process terminatedGate closed (no
LAB_GO):text /* TID 0x103 */ 4 ms getenv("LAB_DELAY") => NULL 4 ms decode(n=29, key=0x5a) 4 ms decode => "http://update.example.com/lab" 4 ms getenv("LAB_GO") => NULL 4 ms puts(s="gate closed") 4 ms | getenv("STDBUF") => NULL 4 ms | getenv("STDBUF0") => NULL 4 ms | getenv("STDBUF1") => NULL 4 ms | getenv("STDBUF2") => NULL 4 ms | getenv("_STDBUF_I") => NULL 4 ms | getenv("_STDBUF_O") => NULL 4 ms | getenv("_STDBUF_E") => NULL Process terminatedThe seven
STDBUFlookups are libc configuring its first stream, not the program.|marks calls made inside another traced function (puts); in the first run they are top level becausefprintf, untraced, opened the stream. The gate-closed run still decodes the URL before checking, so hooking the internal function recovers the indicator even when the sample refuses to act. One spawn attempt failed withunexpectedly timed out while initializing suspended process; a retry succeeded. -
Attach instead of spawning. Start the program with a delay, then attach by name:
bash LAB_GO=1 LAB_DELAY=8 ./tracelab & ./venv/bin/frida-trace -n tracelab -i getenv -i fopen -i puts -i 'tracelab!decode'text Attaching... Started tracing 4 functions. Web UI available at http://localhost:62747/ /* TID 0x103 */ 6267 ms decode(n=29, key=0x5a) 6267 ms decode => "http://update.example.com/lab" 6267 ms getenv("LAB_GO") => "1" 6267 ms fopen(path="tracelab.out", mode="w") ... 6268 ms puts(s="wrote tracelab.out") Process terminatedgetenv("LAB_DELAY")is missing: it ran before Frida attached. Attaching worked becausetracelabis our own ad-hoc signed binary; a protected system binary refuses even spawn mode:text Spawning `/bin/echo`... Failed to attach: unable to access process with pid 92877 from the current user accountWith System Integrity Protection on, Apple's platform binaries cannot be instrumented; Windows and Linux lab VMs have no such restriction for processes you own.
Part C: the same trace on Windows
- frida-trace in the VM. Install
frida-toolsin a virtual environment in the Windows VM, setLAB_GO=1in a command prompt, and runfrida-trace -f C:\lab\dbglab.exe -i "KERNEL32.DLL!OutputDebugStringA" -i "KERNEL32.DLL!GetEnvironmentVariableA". Qualify the module: manykernel32exports have aKernelBasetwin, and an unqualified-ihooks every module that exports the name, so one call can be logged twice. Edit the generatedOutputDebugStringA.jsto logargs[0].readUtf8String(). - API Monitor in the VM. Start API Monitor x64, search the API Filter
for
OutputDebugStringAandGetEnvironmentVariableAand tick them, then Monitor New Process ondbglab.exe. Run it once withLAB_GOunset and once set, and compare the two call lists and their Call Stack panes with the relay trace from Part A. - Cross-check from the kernel side. In the same run, have Procmon record
dbglab.exe. Which of the calls in the API Monitor trace produce a Procmon event, and which never do?
Questions to answer: Why does GetEnvironmentVariableA appear in the
relay and Frida traces but could never appear in an strace or ETW kernel
trace? In Part A, which filter removed more noise: the thread or the return
address, and which one would fail on a sample that injects into another
process? If dbglab.exe resolved OutputDebugStringA with API hashing, which
of the tools in this lab would still log the call? What would a trace of
tracelab look like if the sample issued its write as a raw system call?
Key takeaways
- Tracing records many calls at full speed; debugging examines one moment in depth. Monitor, then trace, then debug what the trace cannot explain.
- In-process tracers (API Monitor, Frida, logging breakpoints) give rich arguments but can be seen and removed; kernel-side sources (Procmon, ETW, strace, eBPF) are coarser but survive user-mode tricks.
- Scope every trace to a question, filter by the caller's module, pick one layer, and log return values.
- Read a trace by following handles, grouping calls into actions, and reading decoded arguments; the API boundary is where encrypted strings appear in cleartext.
- Win32 calls expand into native
ntdllcalls, and some never reach the kernel. A native call made directly from the sample's code is a finding. - Direct and indirect syscalls, unhooking, hook detection and late attachment defeat in-process tracing; pair every hook-based trace with a kernel-side record of the same run.