Leçon 5.5 · Windows pour analystes· 45 min
User Mode, Kernel Mode and the Native API
The ring boundary, ntdll's syscall stubs and WOW64 translation — why malware that talks to the native API bypasses user-mode hooks, and where EDR telemetry now watches instead.
Cette leçon n’est disponible qu’en anglais pour le moment.
Objectifs
- Explain the ring3/ring0 boundary and what the syscall instruction actually does
- Read an ntdll export's disassembly and recognise the syscall-stub pattern
- Explain why calling ntdll or issuing syscalls directly bypasses hooks placed higher in the stack, and why security products responded by watching the boundary itself
- Describe what WOW64 does to a 32-bit process on 64-bit Windows and why that changes what an analyst should expect to observe
- Recognise a kernel-mode rootkit and a BYOVD driver load as the boundary case this lesson does not otherwise cover
Mutexes, Events and Inter-Process Communication closed out the kernel-object namespace: named things a process asks the kernel to create and track. This lesson asks the question underneath that one — how does a user-mode process ask the kernel for anything at all, and what does it mean for an analyst that a sample can ask in more than one way?
Every Windows API you have used in this module eventually turns into a request the kernel serves. Most of the time it does not matter how that request is phrased. It starts to matter the moment a sample chooses an unusual phrasing on purpose — and that choice is one of the more reliable "this author is trying not to be seen" signals you will read.
The boundary itself
x86-64 CPUs run code at different privilege levels, called rings. User-mode
code — everything you have analysed so far, including kernel32.dll — runs at
ring 3, the least privileged level, and cannot directly touch hardware,
other processes' memory, or the kernel's own data structures. The kernel runs
at ring 0, where all of that is allowed. Crossing from one to the other is
not a normal function call; it is a controlled transition the CPU itself
enforces, so that user-mode code cannot forge its way into kernel mode by
jumping to an arbitrary address.
The syscall instruction is that transition on modern
x86-64 Windows (the 32-bit predecessor used sysenter, and read the
system call glossary entry for the general idea
across operating systems). syscall saves the return address and flags,
switches to ring 0, and jumps to a fixed kernel entry point chosen ahead of
time — it cannot land anywhere else. That entry point looks at a number the
caller placed in eax and uses it to index a table of kernel functions, then
runs the one selected, then returns to ring 3 exactly where it left off. The
number is the system service number (SSN), and the table it indexes is the
System Service Descriptor Table (SSDT) — internal to the kernel, not
something user-mode code can read or write directly.
ntdll: the thin wrapper
No application calls syscall by hand under normal circumstances, because the
number a given operation needs is not documented, not part of any stable ABI,
and changes between Windows versions and even between builds. Instead,
ntdll.dll ships one small function per kernel operation — NtCreateFile,
NtWriteVirtualMemory, NtQueryInformationProcess, and hundreds more, all
named Nt* (their aliases exported as Zw* behave identically from user
mode) — and each one is little more than a wrapper around a fixed SSN. The
hooking and rootkits lesson already named the
shape you get when you disassemble one on x64:
mov r10, rcx ; the [calling convention](/en/assembly/calling-conventions) puts
; the first argument in rcx, but syscall clobbers rcx — copy it first
mov eax, <SSN> ; this function's system service number
syscall
retEverything above ntdll — kernel32.dll, advapi32.dll, user32.dll, and
the whole layered API from The Windows API for
Analysts — is written in terms of these
Nt* calls. kernel32!CreateFileW validates and translates its friendlier
parameters and eventually calls ntdll!NtCreateFile. ntdll!NtCreateFile
loads its SSN and executes syscall. Nothing below ntdll is API in the
usual sense at all: it is a numbered table inside the kernel, and the numbers
are an implementation detail Microsoft has never promised to keep stable.
Public reference tables such as j00ru's syscall list exist precisely because
this fragility makes memorising a number for one Windows build useless on the
next.
Why this matters for hooking
Hooking and User-Mode Rootkits showed that
inline and IAT hooks are placed on named functions — most often in ntdll
itself, since that is the one place every higher-level API funnels through
regardless of which DLL the caller used. That placement is also the hook's
limit: a hook on ntdll!NtWriteVirtualMemory only fires for code that calls
that exported function. It has nothing to inspect if a sample instead builds
its own three-instruction stub — mov r10, rcx / mov eax, <SSN> / syscall
— and executes syscall directly, or jumps into the tail of the genuine
ntdll stub partway past its own hooked prologue. Either approach reaches the
kernel through the exact same numbered table entry a hooked call would have
used, without ever executing the bytes the hook patched.
This is why malware that resolves its own SSNs, or copies a clean stub from
disk before calling it, is worth flagging even before you know what the call
does: ordinary programs have no reason to avoid ntdll's own exports. The
same idea extends to indirect syscalls, where a sample still jumps through
a genuine, unmodified ntdll stub rather than issuing syscall itself — a
smaller deviation from normal, aimed specifically at tools that expect the
call to originate inside ntdll's own code range rather than at whether a
hook fired at all.
| Layer a defender might watch | What it sees when a sample calls kernel32!WriteFile | What it sees when a sample uses a hand-built syscall stub |
|---|---|---|
A hook on the kernel32 export | The call, in full | Nothing — never executed |
A hook on the ntdll export | The call, in full | Nothing — the real export was never entered |
| The kernel's own syscall dispatch | The call, in full | The call, in full — the SSN and arguments are identical either way |
| ETW / kernel-mode telemetry | The call, in full | The call, in full |
That bottom row is the point. Nothing about crossing into the kernel changes: the same SSN, the same arguments, the same kernel routine runs. A sample can choose how it asks, but it cannot choose whether the kernel notices — which is exactly why detection has moved down the stack to meet it.
WOW64: a second translation layer
A 32-bit process running on 64-bit Windows adds one more hop. WOW64
("Windows on Windows 64") lets a 32-bit image run essentially unmodified by
giving it 32-bit copies of kernel32.dll, ntdll.dll and the rest, and
quietly translating each call into a 64-bit kernel request behind the scenes.
Concretely, the 32-bit ntdll.dll a WOW64 process loads does not execute
syscall itself; its stubs instead call into wow64cpu.dll, which switches
CPU mode, marshals the 32-bit arguments into 64-bit form, and hands off to the
real 64-bit ntdll.dll to make the actual transition.
Two consequences follow for an analyst:
- A WOW64 process has two
ntdllimages loaded — the 32-bit one the sample's own code calls into, and the 64-bit one doing the real work underneath. Tooling that only walks the 32-bit view of a process (some older or narrowly-scoped utilities) can miss what the 64-bit side is doing; prefer analysis tools that are themselves 64-bit, or that are explicitly WOW64-aware, when the target is a 32-bit sample on a 64-bit host. - 32-bit and 64-bit builds of the same family can behave differently on the same OS, purely because one goes through this extra translation and one does not. Note which bitness you analysed, and do not assume a finding on one transfers to the other without checking.
The native API and the kernel object namespace
The Nt*/Zw* functions in ntdll are collectively the native API —
lower-level, less documented, and closer to what the kernel actually
implements than the Win32 API built on top of it. You have already used part
of it without the name: the kernel object namespace from Mutexes, Events and
Inter-Process Communication (\BaseNamedObjects,
session directories, Global\/Local\ prefixes) is native-API territory —
NtCreateMutant, NtOpenEvent and friends operate directly on that
namespace, and the friendlier CreateMutexW you read about there is simply a
Win32 wrapper around it.
The native API is also where a specific injection primitive lives:
KernelCallbackTable injection
patches an entry in a GUI process's PEB — a table user32.dll uses to call
back into user mode when the kernel delivers a window message — so that a
routine window event ends up executing attacker code, with no
CreateRemoteThread call for a defender to see. It sits at the native-API
layer for the same reason direct syscalls do: the PEB and its
KernelCallbackTable pointer are native-API structures, one level below where
most tooling instruments.
The same native/PEB-level access underlies PEB-walk API resolution and
other dynamic import
resolution tricks: rather than
importing GetProcAddress and LoadLibrary normally (which shows up plainly
in the import table), a sample walks the loaded-module list off the PEB itself
and resolves exports by hand. It is a different motive from direct syscalls —
hiding capability from static analysis rather than calls from a runtime
hook — but the same instinct: go one layer lower than the layer being watched.
See PEB walk API resolution and
The Windows API for Analysts for the
normal GetProcAddress path this bypasses.
Detection has moved to the boundary
Because a hand-built syscall stub is indistinguishable from a genuine one by the time it reaches the kernel, modern endpoint security stopped relying solely on user-mode hooks and instead watches nearer to, or inside, the boundary itself:
- ETW (Event Tracing for Windows) providers built into the kernel can
report activity — process, image-load, network and, through the dedicated
ETW Threat Intelligence (ETWTI) provider, security-relevant operations
such as remote memory allocation/protection changes and remote thread
creation — from inside the kernel, where a user-mode hook bypass has no
effect. ETWTI is what commercial EDR sensors typically consume for exactly
this reason: it reports the same call whether it arrived via
kernel32,ntdll, or a hand-built syscall stub. - Sysmon registers kernel-mode callbacks and a minifilter to observe process creation, image loads, and file/registry activity from the kernel side rather than by hooking user-mode exports — which is why its event log keeps recording even against samples built specifically to dodge user-mode hooking.
- Some EDR products go further and instrument the syscall dispatch path itself (kernel-mode hooks or, more recently, hardware-assisted tracing), explicitly to close the gap direct and indirect syscalls open against purely user-mode instrumentation.
The practical upshot for triage: a sample using direct/indirect syscalls or a PEB walk is trying to evade a specific class of tooling (user-mode API hooks and static import tables), not "detection" in general. Say so precisely in a report — which layer was evaded, and which layer still saw it.
Kernel-mode rootkits and the far side of the boundary
Everything above is still user mode reaching the kernel through the front door, however cleverly phrased. A kernel-mode rootkit is different in kind: a driver, loaded into ring 0 itself, that can alter the very structures and callbacks the detection surfaces above depend on. That is out of scope for this lesson — it is a large topic on its own — but two things are worth recognising when you meet it in the wild:
- A driver load is a loud event by construction: Windows requires kernel drivers to be signed, so a rootkit needs a stolen, leaked, or otherwise abused signing certificate, or must run on a system with driver-signature enforcement weakened.
- BYOVD ("bring your own vulnerable driver") is the dominant real-world pattern: rather than write and sign a malicious driver, an intruder loads a legitimate, signed driver known to contain a vulnerability, and uses that vulnerability to get arbitrary kernel-mode execution. Recognising this pattern statically means recognising the driver, not the exploit — unusual signed drivers being loaded, especially ones with known CVEs and no business reason to be present, is itself the detection signal security products and hunt teams key on.
Lab: watching the boundary from both sides
This lab is entirely observational. You will not write, compile, or execute any code that issues a syscall or resolves one by hand — you will disassemble an existing system file to see the pattern, and separately watch a perfectly ordinary program from the outside.
-
Disassemble a real
ntdllexport. On a Windows machine or VM, copyC:\Windows\System32\ntdll.dllsomewhere you can work with it (or use a 64-bitntdll.dllyou already have from prior lessons' analysis VM). Withpefileand Capstone in a Python virtual environment:python # ntdll_stub.py: show the syscall stub for one named ntdll export import sys import pefile from capstone import Cs, CS_ARCH_X86, CS_MODE_64 pe = pefile.PE(sys.argv[1]) base = pe.OPTIONAL_HEADER.ImageBase image = pe.get_memory_mapped_image() name = sys.argv[2].encode() export = next(e for e in pe.DIRECTORY_ENTRY_EXPORT.symbols if e.name == name) rva = export.address print(f"{name.decode()} at RVA {rva:#x}") md = Cs(CS_ARCH_X86, CS_MODE_64) for i in md.disasm(image[rva:rva + 32], base + rva, 4): print(f" {i.address:#x} {i.bytes.hex(' '):<20} {i.mnemonic} {i.op_str}")bash ./venv/bin/python ntdll_stub.py ntdll.dll NtWriteVirtualMemoryYou should see the same four-instruction shape as the excerpt earlier in this lesson —
mov r10, rcx,mov eax, <SSN>,syscall,ret— with a different SSN for each export you try. Run it again againstNtCreateFile,NtQueryInformationProcessandNtAllocateVirtualMemoryand note that only the number changes; the shape does not. -
Confirm the same pattern with
objdumpas a cross-check:bash objdump -d --start-address=0x<base+rva> --stop-address=0x<base+rva+16> ntdll.dll(Adjust the address range to the RVA
ntdll_stub.pyprinted, added tontdll.dll's image base, or simply grep the full disassembly for the export's symbol.) -
Build one ordinary program. Compile a harmless file-writing program with
mingw-w64— it calls only normal Win32 functions, nothing native or syscall-level:c /* filewrite.c: write one line to a file, using ordinary Win32 calls only */ #include <windows.h> int main(void) { HANDLE h = CreateFileW(L"lab_output.txt", GENERIC_WRITE, 0, NULL, CREATE_ALWAYS, FILE_ATTRIBUTE_NORMAL, NULL); if (h == INVALID_HANDLE_VALUE) return 1; const char msg[] = "user-mode/kernel-mode lab\r\n"; DWORD written; WriteFile(h, msg, sizeof(msg) - 1, &written, NULL); CloseHandle(h); return 0; }bash x86_64-w64-mingw32-gcc -O1 -s -o filewrite.exe filewrite.c -
Watch it from the Win32 side. In your Windows analysis VM (see Building a Safe Analysis Lab), run
filewrite.exeunder Process Monitor with a filter on its process name. You will seeCreateFileandWriteFileevents reported exactly as the source calls them — Process Monitor works by instrumenting well above the syscall boundary, which is why it can present friendly Win32 names at all. Enable Sysmon as well and confirm Event ID 1 (process creation) is logged forfilewrite.exe's launch. -
Locate the same calls in
ntdllterms. Withpefile, listfilewrite.exe's imports and confirm it importskernel32!CreateFileWandkernel32!WriteFile— never anNt*name directly, because the compiler linked against the normal Win32 import libraries. Then repeat step 1 againstNtCreateFileandNtWriteFileinntdll.dllfrom the same VM, and note that these are the two native-API functionskernel32's wrappers call into on your behalf, underneath everything Process Monitor showed you.
Questions to answer: If filewrite.exe had been built to call
NtCreateFile/NtWriteFile directly instead of the Win32 wrappers, which
step above would still have shown its activity, and which would not? Why does
the SSN for the same named export differ between Windows versions, and what
does that imply for a detection rule that hard-codes a number instead of a
name? A WOW64 process's 32-bit ntdll.dll does not itself execute syscall
— which file does, and what would you need to inspect to see the real
transition for a 32-bit sample?
Key takeaways
- User mode (ring 3) reaches the kernel (ring 0) only through the
syscallinstruction, which the CPU restricts to a single fixed entry point indexed by a system service number (SSN);ntdll.dllexports one thin wrapper per SSN, and every higher-level API eventually calls through it. - SSNs are undocumented and version-dependent, which is exactly why nothing
above
ntdllis supposed to depend on them directly — and exactly why a sample that does is worth flagging. - A hand-built or indirect syscall reaches the kernel through the identical code path a hooked, named call would have used; it evades hooks placed on named functions, not the kernel's own visibility, which is why ETW and kernel-mode telemetry (Sysmon, EDR syscall instrumentation) still see it.
- WOW64 adds a translation hop for 32-bit processes on 64-bit Windows,
through a second, 64-bit
ntdll.dllmost user-mode-only tooling never looks at. - Kernel-mode rootkits and BYOVD driver abuse operate below this entire boundary; recognise the loaded driver and its signature as the detection signal, since the kernel-mode payload itself is out of user-mode tooling's reach.