Skip to content

Leçon 3.3 · Triage statique· 35 min

Reading Capabilities from Imports

Turn a PE import table into a capability hypothesis: group APIs by behaviour, spot suspicious combinations, and use capa to confirm with code evidence.

Cette leçon n’est disponible qu’en anglais pour le moment.

Objectifs

  • Group Windows API imports into behaviour categories and form a capability hypothesis
  • Recognise high-signal API combinations that defenders use as detection cues
  • Explain what a very small import table implies about a sample
  • Run capa, read its ATT&CK and MBC output, and separate runtime noise from real findings

In Imports, Exports and the IAT you learned how the import table is built and how the loader fills it. This lesson is about reading it. A Windows program cannot touch files, the registry, the network or other processes without asking the operating system, and those requests go through documented APIs. So the list of APIs a binary imports is, in effect, a list of things it is able to do — a capability hypothesis you can form in minutes, before opening a disassembler.

From imports to a hypothesis

Reading imports well means grouping them. Individual functions are rarely alarming; categories and combinations are. The table below covers the groups you will use most in triage.

BehaviourTypical APIsWhat it may indicate
File systemCreateFileW, WriteFile, DeleteFileW, MoveFileExW, FindFirstFileW, GetTempPathWDropping files, staging, file enumeration (ransomware, stealers)
RegistryRegCreateKeyExW, RegSetValueExW, RegQueryValueExW, RegDeleteKeyWConfiguration storage, persistence (Run keys), system discovery
Process creationCreateProcessW, ShellExecuteExW, WinExecLaunching payloads or commands
Process discoveryCreateToolhelp32Snapshot, Process32FirstW, EnumProcessesLooking for security tools, analysis tools or a target process
Memory in other processesOpenProcess, VirtualAllocEx, WriteProcessMemory, CreateRemoteThread, QueueUserAPC, SetThreadContextCode injection (see below)
NetworkingWinINet (InternetOpenW, InternetOpenUrlW, HttpSendRequestW), WinHTTP (WinHttpConnect), Winsock (socket, connect, send, recv), URLDownloadToFileWC2 communication, downloading further stages, exfiltration
CryptographyCryptAcquireContextW, CryptEncrypt, BCryptEncrypt, BCryptGenRandomEncrypting config or traffic; bulk use suggests ransomware
ServicesOpenSCManagerW, CreateServiceW, StartServiceWPersistence or privilege use via services
Anti-debuggingIsDebuggerPresent, CheckRemoteDebuggerPresent, NtQueryInformationProcess, OutputDebugStringWEnvironment checks that alter behaviour under analysis
Input and screenSetWindowsHookExW, GetAsyncKeyState, GetClipboardData, BitBltKeylogging, clipboard theft, screenshots
Own resourcesFindResourceW, LoadResource, LockResource, SizeofResourceEmbedded data or payload (see Resources and Overlays)
Dynamic resolutionLoadLibraryW, GetProcAddress, LdrGetProcedureAddressImports hidden from this table

Write the hypothesis down as short statements you intend to confirm or refute later: "can read and write registry values; can make HTTP requests; loads its own resource". That list drives what you look for in dynamic analysis and where you set breakpoints.

Combinations carry the signal

Detection engineers care most about sets of imports that only make sense together. The best known is the classic remote-injection sequence:

text
  OpenProcess ─► VirtualAllocEx ─► WriteProcessMemory ─► CreateRemoteThread
  (get a handle)  (memory in target)  (copy bytes there)    (run them)

Each API alone has legitimate uses — debuggers and installers call several of them. The full chain in a small, unsigned binary is a strong indicator of process injection, which is why EDR products watch for it at runtime and why capa has rules for it. The CreateRemoteThread DLL injection page describes the technique from the defender's side.

Other combinations worth a second look:

  • FindResource + LoadResource + VirtualAlloc + VirtualProtect — unpacking or decrypting an embedded payload into executable memory.
  • CreateToolhelp32Snapshot + Process32Next + TerminateProcess — hunting and killing security products.
  • CryptEncrypt or BCryptEncrypt + FindFirstFile/FindNextFile + MoveFileEx — encrypt-and-rename loops typical of ransomware.
  • RegSetValueEx together with a string containing Software\Microsoft\Windows\CurrentVersion\Run — Run-key persistence.

Tip: Always read imports together with strings (see Strings and Obfuscated Strings). An import says how; a string such as a registry path, URL or process name often says what.

Reading API names precisely

Windows API names follow conventions that change their meaning slightly:

  • A and W suffixes. CreateFileA takes ANSI strings, CreateFileW takes UTF-16. Modern code mostly imports W; many hand-written samples import A. The suffix also tells you which string encoding to search for.
  • Ex suffixes. An extended version with more parameters — VirtualAllocEx is VirtualAlloc for another process, which is exactly why it matters.
  • Nt / Zw functions from ntdll.dll. These are the native API underneath the documented Win32 layer. Legitimate programs rarely import them directly. Direct ntdll imports such as NtAllocateVirtualMemory, NtWriteVirtualMemory or NtQueryInformationProcess suggest code that is trying to sit below user-mode hooks, or to query process information for anti-debugging (see IsDebuggerPresent and CheckRemoteDebuggerPresent).
  • Ordinals. Some DLLs are commonly imported by number. ws2_32.dll is the classic case: ordinal 23 is socket, ordinal 115 is WSAStartup. pefile and most GUI tools translate well-known ordinals back to names.

When the table is almost empty

A normal compiled program imports dozens to hundreds of functions. A table with a handful of entries is itself a finding. For example, this is the entire KERNEL32.DLL import list of a small program after packing it with UPX:

text
KERNEL32.DLL   LoadLibraryA, ExitProcess, GetProcAddress, VirtualProtect

With LoadLibrary and GetProcAddress, code can resolve any API it wants at runtime, so the real capability list has moved out of the table and into code or encrypted data. The usual explanations are:

  • Packing. The unpacking stub needs only a few APIs; the original imports are rebuilt in memory. See Detecting Packers and Entropy and UPX packing.
  • Dynamic resolution. The author calls GetProcAddress with names built at runtime, often as stack strings — see dynamic import resolution.
  • API hashing. Even GetProcAddress disappears: the code walks the export tables of loaded DLLs and compares name hashes. See API hashing.

In all three cases, static import analysis has reached its limit and the next step is to unpack, or to observe the resolved APIs dynamically.

Imports show possibility, not execution

Keep two caveats in mind:

  1. An import is not a call path. A function may be imported but never reached, or reached only under conditions that never occur. Imports justify a hypothesis, not a verdict.
  2. Libraries and runtimes add noise. Statically linked libraries and the C runtime import functions your target code never asked for. You will see a concrete case in the lab: MinGW's startup code imports VirtualProtect and changes memory protections, which makes a harmless program look like it manipulates executable memory.

capa: capabilities with evidence

capa (from Mandiant's FLARE team) automates this reasoning and goes further. It disassembles the program and matches a large community rule set against features it extracts — API calls, strings, constants, byte sequences, and the structure of the functions that use them. Rules are scoped (to a basic block, a function, or the whole file), so a rule like "create HTTP request" can require that InternetOpen and HttpSendRequest are called from the same function, which is much stronger evidence than both being present in the import table.

Each match is labelled with:

  • a capability name and namespace, such as communication/http/client;
  • MITRE ATT&CK techniques, useful for reports and for correlating with detections;
  • MBC (Malware Behavior Catalog) behaviours, a malware-specific taxonomy that complements ATT&CK.

Run it as capa sample.exe for a summary, capa -v to see the addresses that matched each rule, and capa -vv for the full evidence tree. The verbose modes tell you exactly which functions to open in your disassembler.

Warning: capa is only as good as what it can see. On a packed sample it will say so and match little besides "packed with …" rules — unpack first, then run it again.

Lab: predict, then verify

This lab uses three small, harmless programs that each exercise one behaviour category. You need mingw-w64, Python 3 with pefile, and capa (the standalone release binary, or pip install flare-capa plus the rules repository).

  1. Create the programs.

    regread.c — reads (and creates if missing) a value under HKCU\Software\LabDemo:

    c
    #include <windows.h>
    #include <stdio.h>
    
    int main(void) {
        HKEY key;
        char buf[64] = {0};
        DWORD size = sizeof(buf), type = 0;
        if (RegCreateKeyExA(HKEY_CURRENT_USER, "Software\\LabDemo", 0, NULL, 0,
                            KEY_READ | KEY_WRITE, NULL, &key, NULL) != ERROR_SUCCESS)
            return 1;
        if (RegQueryValueExA(key, "Greeting", NULL, &type, (BYTE *)buf, &size) != ERROR_SUCCESS) {
            const char *def = "hello from the lab";
            RegSetValueExA(key, "Greeting", 0, REG_SZ, (const BYTE *)def, (DWORD)strlen(def) + 1);
            strcpy(buf, def);
        }
        printf("Greeting = %s\n", buf);
        RegCloseKey(key);
        return 0;
    }

    httpget.c — downloads http://example.com/ with WinINet and prints the byte count:

    c
    #include <windows.h>
    #include <wininet.h>
    #include <stdio.h>
    
    int main(void) {
        char buf[4096];
        DWORD got, total = 0;
        HINTERNET net = InternetOpenA("LabDemo/1.0", INTERNET_OPEN_TYPE_PRECONFIG, NULL, NULL, 0);
        if (!net) return 1;
        HINTERNET url = InternetOpenUrlA(net, "http://example.com/", NULL, 0,
                                         INTERNET_FLAG_RELOAD | INTERNET_FLAG_NO_CACHE_WRITE, 0);
        if (url) {
            while (InternetReadFile(url, buf, sizeof(buf), &got) && got) total += got;
            InternetCloseHandle(url);
        }
        InternetCloseHandle(net);
        printf("read %lu bytes\n", total);
        return 0;
    }

    proclist.c — lists running processes:

    c
    #include <windows.h>
    #include <tlhelp32.h>
    #include <stdio.h>
    
    int main(void) {
        HANDLE snap = CreateToolhelp32Snapshot(TH32CS_SNAPPROCESS, 0);
        if (snap == INVALID_HANDLE_VALUE) return 1;
        PROCESSENTRY32 pe = { .dwSize = sizeof(pe) };
        for (BOOL ok = Process32First(snap, &pe); ok; ok = Process32Next(snap, &pe))
            printf("%6lu  %s\n", pe.th32ProcessID, pe.szExeFile);
        CloseHandle(snap);
        return 0;
    }
  2. Build them:

    bash
    x86_64-w64-mingw32-gcc -s -o regread.exe regread.c -ladvapi32
    x86_64-w64-mingw32-gcc -s -o httpget.exe httpget.c -lwininet
    x86_64-w64-mingw32-gcc -s -o proclist.exe proclist.c
  3. Predict before you look at any tool output. Swap the files with a lab partner (or wait a day), then list only the imports and write one line of capability hypothesis for each binary. This script groups imports into the categories from the table above:

    python
    # capimports.py — group a PE's imports into capability buckets
    import re, sys
    import pefile
    
    BUCKETS = {
        "file system":  r"^(CreateFile|ReadFile|WriteFile|DeleteFile|MoveFile|CopyFile|FindFirstFile|FindNextFile)",
        "registry":     r"^Reg(Open|Create|Set|Query|Delete|Enum)",
        "process":      r"^(CreateProcess|OpenProcess|TerminateProcess|CreateToolhelp32Snapshot|Process32|ShellExecute|WinExec)",
        "memory":       r"^(VirtualAlloc|VirtualProtect|WriteProcessMemory|ReadProcessMemory|CreateRemoteThread|QueueUserAPC|SetThreadContext)",
        "network":      r"^(Internet|Http|WinHttp|URLDownload|WSA|socket$|connect$|send$|recv$)",
        "crypto":       r"^(Crypt|BCrypt)",
        "anti-debug":   r"^(IsDebuggerPresent|CheckRemoteDebuggerPresent|NtQueryInformationProcess)",
        "resources":    r"^(FindResource|LoadResource|LockResource|SizeofResource)",
        "dynamic":      r"^(LoadLibrary|GetProcAddress)",
    }
    
    pe = pefile.PE(sys.argv[1])
    hits, total = {}, 0
    for desc in getattr(pe, "DIRECTORY_ENTRY_IMPORT", []):
        dll = desc.dll.decode().lower()
        for imp in desc.imports:
            total += 1
            name = imp.name.decode() if imp.name else f"ord{imp.ordinal}"
            for bucket, rx in BUCKETS.items():
                if re.match(rx, name):
                    hits.setdefault(bucket, []).append(f"{name} ({dll})")
    print(f"{sys.argv[1]}: {total} imports")
    for bucket, names in hits.items():
        print(f"  {bucket:13} {', '.join(names)}")

    With a recent MinGW build, the output looks like this:

    text
    regread.exe: 45 imports
      registry      RegCreateKeyExA (advapi32.dll), RegQueryValueExA (advapi32.dll), RegSetValueExA (advapi32.dll)
      memory        VirtualProtect (kernel32.dll)
    httpget.exe: 45 imports
      memory        VirtualProtect (kernel32.dll)
      network       InternetCloseHandle (wininet.dll), InternetOpenA (wininet.dll), InternetOpenUrlA (wininet.dll), InternetReadFile (wininet.dll)
    proclist.exe: 45 imports
      process       CreateToolhelp32Snapshot (kernel32.dll), Process32First (kernel32.dll), Process32Next (kernel32.dll)
      memory        VirtualProtect (kernel32.dll)

    Exact counts depend on your MinGW version and C runtime.

  4. Now run capa on each binary and compare with your hypotheses:

    bash
    capa httpget.exe

    For httpget.exe, the capability table includes entries such as:

    text
    receive data                            communication
    create HTTP request                      communication/http/client
    allocate or change RWX memory           host-interaction/process/inject
    terminate process                       host-interaction/process/terminate
    enumerate PE sections                   load-code/pe
    parse PE header                         load-code/pe

    regread.exe adds "query or enumerate registry value" and "set registry value" (ATT&CK T1012, Query Registry), and proclist.exe adds "enumerate processes".

  5. Investigate the surprise. All three programs — and a plain "hello world" — match "allocate or change RWX memory" under the inject namespace, plus "parse PE header". Run capa -v httpget.exe, note the function addresses for that rule, and confirm in Ghidra or any disassembler that they belong to the MinGW runtime startup code (which calls VirtualProtect with PAGE_EXECUTE_READWRITE, 0x40), not to main.

  6. Pack one binary with UPX and repeat steps 3 and 4:

    bash
    upx -o httpget-upx.exe httpget.exe
    python3 capimports.py httpget-upx.exe
    capa httpget-upx.exe

    The WinINet imports vanish from the table, and capa warns that the file appears packed.

Questions to answer: Which of your one-line hypotheses were right, and which did capa refine with code evidence? Why would a rule scoped to a single function produce fewer false positives than an import-table check? How would you word the RWX finding in a report so a reader does not mistake runtime noise for injection?

Key takeaways

  • Imports describe what a program can ask Windows to do; group them into behaviour categories and write a short capability hypothesis.
  • Combinations such as OpenProcess → VirtualAllocEx → WriteProcessMemory → CreateRemoteThread carry far more signal than any single API.
  • A/W, Ex, Nt/Zw and ordinals all change what an import tells you.
  • A tiny import table with LoadLibrary/GetProcAddress means the real imports are hidden — by packing, dynamic resolution or API hashing.
  • capa confirms capabilities with code-level evidence and maps them to ATT&CK and MBC, but runtime and library code produce matches too: check where each match lives before you report it.