Skip to content

Lesson 5.5 · Windows Internals for Analysts· 45 min

User Mode, Kernel Mode and the Native API

The ring boundary, ntdll's syscall stubs and WOW64 translation — why malware that talks to the native API bypasses user-mode hooks, and where EDR telemetry now watches instead.

Objectives

  • Explain the ring3/ring0 boundary and what the syscall instruction actually does
  • Read an ntdll export's disassembly and recognise the syscall-stub pattern
  • Explain why calling ntdll or issuing syscalls directly bypasses hooks placed higher in the stack, and why security products responded by watching the boundary itself
  • Describe what WOW64 does to a 32-bit process on 64-bit Windows and why that changes what an analyst should expect to observe
  • Recognise a kernel-mode rootkit and a BYOVD driver load as the boundary case this lesson does not otherwise cover

Mutexes, Events and Inter-Process Communication closed out the kernel-object namespace: named things a process asks the kernel to create and track. This lesson asks the question underneath that one — how does a user-mode process ask the kernel for anything at all, and what does it mean for an analyst that a sample can ask in more than one way?

Every Windows API you have used in this module eventually turns into a request the kernel serves. Most of the time it does not matter how that request is phrased. It starts to matter the moment a sample chooses an unusual phrasing on purpose — and that choice is one of the more reliable "this author is trying not to be seen" signals you will read.

The boundary itself

x86-64 CPUs run code at different privilege levels, called rings. User-mode code — everything you have analysed so far, including kernel32.dll — runs at ring 3, the least privileged level, and cannot directly touch hardware, other processes' memory, or the kernel's own data structures. The kernel runs at ring 0, where all of that is allowed. Crossing from one to the other is not a normal function call; it is a controlled transition the CPU itself enforces, so that user-mode code cannot forge its way into kernel mode by jumping to an arbitrary address.

The syscall instruction is that transition on modern x86-64 Windows (the 32-bit predecessor used sysenter, and read the system call glossary entry for the general idea across operating systems). syscall saves the return address and flags, switches to ring 0, and jumps to a fixed kernel entry point chosen ahead of time — it cannot land anywhere else. That entry point looks at a number the caller placed in eax and uses it to index a table of kernel functions, then runs the one selected, then returns to ring 3 exactly where it left off. The number is the system service number (SSN), and the table it indexes is the System Service Descriptor Table (SSDT) — internal to the kernel, not something user-mode code can read or write directly.

ntdll: the thin wrapper

No application calls syscall by hand under normal circumstances, because the number a given operation needs is not documented, not part of any stable ABI, and changes between Windows versions and even between builds. Instead, ntdll.dll ships one small function per kernel operation — NtCreateFile, NtWriteVirtualMemory, NtQueryInformationProcess, and hundreds more, all named Nt* (their aliases exported as Zw* behave identically from user mode) — and each one is little more than a wrapper around a fixed SSN. The hooking and rootkits lesson already named the shape you get when you disassemble one on x64:

text
mov r10, rcx        ; the [calling convention](/en/assembly/calling-conventions) puts
                     ; the first argument in rcx, but syscall clobbers rcx — copy it first
mov eax, <SSN>       ; this function's system service number
syscall
ret

Everything above ntdll — kernel32.dll, advapi32.dll, user32.dll, and the whole layered API from The Windows API for Analysts — is written in terms of these Nt* calls. kernel32!CreateFileW validates and translates its friendlier parameters and eventually calls ntdll!NtCreateFile. ntdll!NtCreateFile loads its SSN and executes syscall. Nothing below ntdll is API in the usual sense at all: it is a numbered table inside the kernel, and the numbers are an implementation detail Microsoft has never promised to keep stable. Public reference tables such as j00ru's syscall list exist precisely because this fragility makes memorising a number for one Windows build useless on the next.

Why this matters for hooking

Hooking and User-Mode Rootkits showed that inline and IAT hooks are placed on named functions — most often in ntdll itself, since that is the one place every higher-level API funnels through regardless of which DLL the caller used. That placement is also the hook's limit: a hook on ntdll!NtWriteVirtualMemory only fires for code that calls that exported function. It has nothing to inspect if a sample instead builds its own three-instruction stub — mov r10, rcx / mov eax, <SSN> / syscall — and executes syscall directly, or jumps into the tail of the genuine ntdll stub partway past its own hooked prologue. Either approach reaches the kernel through the exact same numbered table entry a hooked call would have used, without ever executing the bytes the hook patched.

This is why malware that resolves its own SSNs, or copies a clean stub from disk before calling it, is worth flagging even before you know what the call does: ordinary programs have no reason to avoid ntdll's own exports. The same idea extends to indirect syscalls, where a sample still jumps through a genuine, unmodified ntdll stub rather than issuing syscall itself — a smaller deviation from normal, aimed specifically at tools that expect the call to originate inside ntdll's own code range rather than at whether a hook fired at all.

Layer a defender might watchWhat it sees when a sample calls kernel32!WriteFileWhat it sees when a sample uses a hand-built syscall stub
A hook on the kernel32 exportThe call, in fullNothing — never executed
A hook on the ntdll exportThe call, in fullNothing — the real export was never entered
The kernel's own syscall dispatchThe call, in fullThe call, in full — the SSN and arguments are identical either way
ETW / kernel-mode telemetryThe call, in fullThe call, in full

That bottom row is the point. Nothing about crossing into the kernel changes: the same SSN, the same arguments, the same kernel routine runs. A sample can choose how it asks, but it cannot choose whether the kernel notices — which is exactly why detection has moved down the stack to meet it.

WOW64: a second translation layer

A 32-bit process running on 64-bit Windows adds one more hop. WOW64 ("Windows on Windows 64") lets a 32-bit image run essentially unmodified by giving it 32-bit copies of kernel32.dll, ntdll.dll and the rest, and quietly translating each call into a 64-bit kernel request behind the scenes. Concretely, the 32-bit ntdll.dll a WOW64 process loads does not execute syscall itself; its stubs instead call into wow64cpu.dll, which switches CPU mode, marshals the 32-bit arguments into 64-bit form, and hands off to the real 64-bit ntdll.dll to make the actual transition.

Two consequences follow for an analyst:

  • A WOW64 process has two ntdll images loaded — the 32-bit one the sample's own code calls into, and the 64-bit one doing the real work underneath. Tooling that only walks the 32-bit view of a process (some older or narrowly-scoped utilities) can miss what the 64-bit side is doing; prefer analysis tools that are themselves 64-bit, or that are explicitly WOW64-aware, when the target is a 32-bit sample on a 64-bit host.
  • 32-bit and 64-bit builds of the same family can behave differently on the same OS, purely because one goes through this extra translation and one does not. Note which bitness you analysed, and do not assume a finding on one transfers to the other without checking.

The native API and the kernel object namespace

The Nt*/Zw* functions in ntdll are collectively the native API — lower-level, less documented, and closer to what the kernel actually implements than the Win32 API built on top of it. You have already used part of it without the name: the kernel object namespace from Mutexes, Events and Inter-Process Communication (\BaseNamedObjects, session directories, Global\/Local\ prefixes) is native-API territory — NtCreateMutant, NtOpenEvent and friends operate directly on that namespace, and the friendlier CreateMutexW you read about there is simply a Win32 wrapper around it.

The native API is also where a specific injection primitive lives: KernelCallbackTable injection patches an entry in a GUI process's PEB — a table user32.dll uses to call back into user mode when the kernel delivers a window message — so that a routine window event ends up executing attacker code, with no CreateRemoteThread call for a defender to see. It sits at the native-API layer for the same reason direct syscalls do: the PEB and its KernelCallbackTable pointer are native-API structures, one level below where most tooling instruments.

The same native/PEB-level access underlies PEB-walk API resolution and other dynamic import resolution tricks: rather than importing GetProcAddress and LoadLibrary normally (which shows up plainly in the import table), a sample walks the loaded-module list off the PEB itself and resolves exports by hand. It is a different motive from direct syscalls — hiding capability from static analysis rather than calls from a runtime hook — but the same instinct: go one layer lower than the layer being watched. See PEB walk API resolution and The Windows API for Analysts for the normal GetProcAddress path this bypasses.

Detection has moved to the boundary

Because a hand-built syscall stub is indistinguishable from a genuine one by the time it reaches the kernel, modern endpoint security stopped relying solely on user-mode hooks and instead watches nearer to, or inside, the boundary itself:

  • ETW (Event Tracing for Windows) providers built into the kernel can report activity — process, image-load, network and, through the dedicated ETW Threat Intelligence (ETWTI) provider, security-relevant operations such as remote memory allocation/protection changes and remote thread creation — from inside the kernel, where a user-mode hook bypass has no effect. ETWTI is what commercial EDR sensors typically consume for exactly this reason: it reports the same call whether it arrived via kernel32, ntdll, or a hand-built syscall stub.
  • Sysmon registers kernel-mode callbacks and a minifilter to observe process creation, image loads, and file/registry activity from the kernel side rather than by hooking user-mode exports — which is why its event log keeps recording even against samples built specifically to dodge user-mode hooking.
  • Some EDR products go further and instrument the syscall dispatch path itself (kernel-mode hooks or, more recently, hardware-assisted tracing), explicitly to close the gap direct and indirect syscalls open against purely user-mode instrumentation.

The practical upshot for triage: a sample using direct/indirect syscalls or a PEB walk is trying to evade a specific class of tooling (user-mode API hooks and static import tables), not "detection" in general. Say so precisely in a report — which layer was evaded, and which layer still saw it.

Kernel-mode rootkits and the far side of the boundary

Everything above is still user mode reaching the kernel through the front door, however cleverly phrased. A kernel-mode rootkit is different in kind: a driver, loaded into ring 0 itself, that can alter the very structures and callbacks the detection surfaces above depend on. That is out of scope for this lesson — it is a large topic on its own — but two things are worth recognising when you meet it in the wild:

  • A driver load is a loud event by construction: Windows requires kernel drivers to be signed, so a rootkit needs a stolen, leaked, or otherwise abused signing certificate, or must run on a system with driver-signature enforcement weakened.
  • BYOVD ("bring your own vulnerable driver") is the dominant real-world pattern: rather than write and sign a malicious driver, an intruder loads a legitimate, signed driver known to contain a vulnerability, and uses that vulnerability to get arbitrary kernel-mode execution. Recognising this pattern statically means recognising the driver, not the exploit — unusual signed drivers being loaded, especially ones with known CVEs and no business reason to be present, is itself the detection signal security products and hunt teams key on.

Lab: watching the boundary from both sides

This lab is entirely observational. You will not write, compile, or execute any code that issues a syscall or resolves one by hand — you will disassemble an existing system file to see the pattern, and separately watch a perfectly ordinary program from the outside.

  1. Disassemble a real ntdll export. On a Windows machine or VM, copy C:\Windows\System32\ntdll.dll somewhere you can work with it (or use a 64-bit ntdll.dll you already have from prior lessons' analysis VM). With pefile and Capstone in a Python virtual environment:

    python
    # ntdll_stub.py: show the syscall stub for one named ntdll export
    import sys
    import pefile
    from capstone import Cs, CS_ARCH_X86, CS_MODE_64
    
    pe = pefile.PE(sys.argv[1])
    base = pe.OPTIONAL_HEADER.ImageBase
    image = pe.get_memory_mapped_image()
    
    name = sys.argv[2].encode()
    export = next(e for e in pe.DIRECTORY_ENTRY_EXPORT.symbols if e.name == name)
    rva = export.address
    print(f"{name.decode()} at RVA {rva:#x}")
    
    md = Cs(CS_ARCH_X86, CS_MODE_64)
    for i in md.disasm(image[rva:rva + 32], base + rva, 4):
        print(f"  {i.address:#x}  {i.bytes.hex(' '):<20} {i.mnemonic} {i.op_str}")
    bash
    ./venv/bin/python ntdll_stub.py ntdll.dll NtWriteVirtualMemory

    You should see the same four-instruction shape as the excerpt earlier in this lesson — mov r10, rcx, mov eax, <SSN>, syscall, ret — with a different SSN for each export you try. Run it again against NtCreateFile, NtQueryInformationProcess and NtAllocateVirtualMemory and note that only the number changes; the shape does not.

  2. Confirm the same pattern with objdump as a cross-check:

    bash
    objdump -d --start-address=0x<base+rva> --stop-address=0x<base+rva+16> ntdll.dll

    (Adjust the address range to the RVA ntdll_stub.py printed, added to ntdll.dll's image base, or simply grep the full disassembly for the export's symbol.)

  3. Build one ordinary program. Compile a harmless file-writing program with mingw-w64 — it calls only normal Win32 functions, nothing native or syscall-level:

    c
    /* filewrite.c: write one line to a file, using ordinary Win32 calls only */
    #include <windows.h>
    
    int main(void) {
        HANDLE h = CreateFileW(L"lab_output.txt", GENERIC_WRITE, 0, NULL,
                                CREATE_ALWAYS, FILE_ATTRIBUTE_NORMAL, NULL);
        if (h == INVALID_HANDLE_VALUE) return 1;
        const char msg[] = "user-mode/kernel-mode lab\r\n";
        DWORD written;
        WriteFile(h, msg, sizeof(msg) - 1, &written, NULL);
        CloseHandle(h);
        return 0;
    }
    bash
    x86_64-w64-mingw32-gcc -O1 -s -o filewrite.exe filewrite.c
  4. Watch it from the Win32 side. In your Windows analysis VM (see Building a Safe Analysis Lab), run filewrite.exe under Process Monitor with a filter on its process name. You will see CreateFile and WriteFile events reported exactly as the source calls them — Process Monitor works by instrumenting well above the syscall boundary, which is why it can present friendly Win32 names at all. Enable Sysmon as well and confirm Event ID 1 (process creation) is logged for filewrite.exe's launch.

  5. Locate the same calls in ntdll terms. With pefile, list filewrite.exe's imports and confirm it imports kernel32!CreateFileW and kernel32!WriteFile — never an Nt* name directly, because the compiler linked against the normal Win32 import libraries. Then repeat step 1 against NtCreateFile and NtWriteFile in ntdll.dll from the same VM, and note that these are the two native-API functions kernel32's wrappers call into on your behalf, underneath everything Process Monitor showed you.

Questions to answer: If filewrite.exe had been built to call NtCreateFile/NtWriteFile directly instead of the Win32 wrappers, which step above would still have shown its activity, and which would not? Why does the SSN for the same named export differ between Windows versions, and what does that imply for a detection rule that hard-codes a number instead of a name? A WOW64 process's 32-bit ntdll.dll does not itself execute syscall — which file does, and what would you need to inspect to see the real transition for a 32-bit sample?

Key takeaways

  • User mode (ring 3) reaches the kernel (ring 0) only through the syscall instruction, which the CPU restricts to a single fixed entry point indexed by a system service number (SSN); ntdll.dll exports one thin wrapper per SSN, and every higher-level API eventually calls through it.
  • SSNs are undocumented and version-dependent, which is exactly why nothing above ntdll is supposed to depend on them directly — and exactly why a sample that does is worth flagging.
  • A hand-built or indirect syscall reaches the kernel through the identical code path a hooked, named call would have used; it evades hooks placed on named functions, not the kernel's own visibility, which is why ETW and kernel-mode telemetry (Sysmon, EDR syscall instrumentation) still see it.
  • WOW64 adds a translation hop for 32-bit processes on 64-bit Windows, through a second, 64-bit ntdll.dll most user-mode-only tooling never looks at.
  • Kernel-mode rootkits and BYOVD driver abuse operate below this entire boundary; recognise the loaded driver and its signature as the detection signal, since the kernel-mode payload itself is out of user-mode tooling's reach.