Skip to content

Leçon 1.3 · Fondations· 30 min

From Source Code to Binary

Follow a C program through preprocessing, compilation, assembly and linking, and learn what symbols, relocations and debug info leave behind.

Cette leçon n’est disponible qu’en anglais pour le moment.

Objectifs

  • Name the four stages of a C build and produce each intermediate file with GCC and MSVC
  • Explain what object files, symbols and relocations are, and what the linker does with them
  • Distinguish static from dynamic linking and debug info from symbols, and predict what stripping removes
  • Use build artefacts such as PDB paths and library code to speed up malware analysis

Every sample you analyse was, at some point, source code on an attacker's machine. A compiler and a linker turned it into the binary on your desk, and each step both destroyed information you would like to have and added information the author never meant to share. Knowing the pipeline tells you what to expect in a binary, what is gone for good, and where the leftovers hide.

The pipeline at a glance

A C build is really four tools run in sequence. Compiler drivers such as gcc and cl hide this by running all of them for you, but each stage can be stopped and inspected:

text
  beacon.c ──► [preprocessor] ──► beacon.i ──► [compiler] ──► beacon.s
   (source)      #include,          (pure C)    optimise,     (assembly
                 #define, #if                    lower to asm   text)
                                                                  │
                                                                  ▼
  beacon ◄── [linker] ◄── beacon.o + other .o + libraries ◄── [assembler]
  (.exe/ELF)  resolve symbols,     (object files: machine code
              apply relocations,    with holes still unfilled)
              lay out sections
StageInput → outputGCCMSVC
Preprocess.c → .igcc -Ecl /P
Compile.i → .s / .asmgcc -Scl /FA (listing)
Assemble.s → .o / .objgcc -c (or as)cl /c (or ml64)
Link.o + libs → executablegcc (runs ld)link.exe

MSVC does not really go through a text assembly file — /FA writes a listing alongside the object file it produces directly — but the conceptual stages are the same.

Stage 1: preprocessing

The preprocessor is a text transformer. It pastes in every #included header, expands #define macros and evaluates #if/#ifdef blocks. Its output is still C, just enormous: a ten-line program that includes <windows.h> preprocesses to hundreds of thousands of lines.

For analysts, the important consequence is that macros and conditional compilation vanish completely. A malware builder that toggles features with #ifdef ENABLE_KEYLOGGER produces binaries in which disabled features simply do not exist — not as dead code, not as strings, not at all. Two samples from the same family can therefore differ substantially without either being obfuscated.

Stage 2: compilation

The compiler proper parses the preprocessed C, type-checks it, optimises it and emits assembly for the target architecture. This is where most of the information loss described in What Is Malware Analysis? happens: local variable names become stack offsets or registers, for loops become compare-and-jump sequences, and small functions may be inlined so that they no longer exist as separate functions at all.

Optimisation level matters enormously. The same function compiled at -O0 and -O2 (or MSVC /Od and /O2) can look like two unrelated pieces of code. Unoptimised code keeps every variable in memory and is easy to map back to source; optimised code is shorter, keeps values in registers, and reorders work. Most release-mode malware is optimised.

The assembly output still contains names — labels for functions, global variables and compiler-generated labels for string literals — because the next stage needs them.

Stage 3: assembly and object files

The assembler translates assembly text into machine code and packages it as an object file: COFF .obj on Windows, ELF relocatable .o on Linux. An object file already has sections (.text for code, .data and .rodata / .rdata for data) but it is not runnable, for one reason: it has holes.

When beacon.c calls puts, the assembler has no idea where puts will live — it is in a different file entirely. Even references to the file's own strings cannot be finalised, because the linker has not yet decided where each section goes. So the assembler writes placeholder bytes (often zeros) and records a note for each hole:

  • A symbol table listing every name the file defines (such as main) and every name it needs but does not define (such as puts).
  • A relocation table: "at offset X in .text, patch in the address of symbol Y, computed in way Z." The "way" is a relocation type such as R_X86_64_PLT32 on ELF or IMAGE_REL_AMD64_REL32 on COFF — both meaning, roughly, "a 32-bit displacement relative to the next instruction."
text
  object file .text, before linking (illustrative offsets)

  offset  bytes                  instruction          relocation record
  ──────  ─────────────────────  ───────────────────  ──────────────────────
  0x0c    48 8d 05 00 00 00 00   lea rax,[rip+????]   → address of string in .rodata
  0x16    e8 00 00 00 00         call ????            → address of puts
                   ▲
                   └── zero placeholders the linker will overwrite

A relocation is thus a promise: "someone will fill this in later." You will meet relocations again in the loaded program, where the Windows loader applies base relocations when an image cannot sit at its preferred address — see How a Binary Is Loaded and Run.

Stage 4: linking

The linker takes all object files plus libraries and produces one executable. It does three main jobs:

  1. Symbol resolution — match every undefined symbol to exactly one definition. A missing definition gives the familiar "unresolved external symbol" (MSVC) or "undefined reference" (GNU ld) error.
  2. Layout — merge same-named sections from all inputs (every .text into one .text, and so on) and assign each a final virtual address.
  3. Relocation — now that addresses are known, patch every hole the assemblers left.

Static versus dynamic linking

Library functions can reach the final binary in two ways:

Static linkingDynamic linking
Library codeCopied into the executableStays in a separate DLL / .so
Binary sizeLargerSmaller
How calls reach the libraryDirect calls to embedded codeThrough the import table: IAT on Windows, GOT/PLT on Linux
What the analyst seesThousands of anonymous library functionsNamed imports like CreateFileW

With dynamic linking, the linker cannot patch in the final address of CreateFileW, because kernel32.dll is loaded at runtime. Instead it writes an import table naming the DLL and function, and a table of pointer slots — the IAT on Windows, the GOT on Linux — that the loader fills in when the program starts. Those names survive even in a stripped binary, which is why imports are one of the first things you look at during triage.

Static linking is common in malware for practical reasons: the binary runs without dependencies. Go and Rust binaries link their runtimes statically by default, and Linux IoT malware is typically statically linked so it runs on any device. The analyst's cost is noise: a stripped, statically linked binary presents thousands of unnamed functions, of which the author wrote perhaps a hundred.

Symbols and debug information

Two different kinds of metadata are easy to confuse:

  • Symbols map names to addresses: "main is at 0x1149." They are what nm lists and what the linker needs.
  • Debug information is far richer: types, local variable names and locations, source file names and line-number mappings. It exists for debuggers.
Linux / ELFWindows / PE (MSVC)
Symbols.symtab (all) and .dynsym (imports/exports) sectionsNot in the image; in the PDB
Debug infoDWARF, in .debug_* sections of the same filePDB, a separate file
Produced bygcc -gcl /Zi + link /DEBUG
Removed bystripNot shipping the .pdb

strip removes .symtab and the DWARF sections, but not .dynsym: the dynamic loader needs those names to resolve imports and exports at runtime. PE images built by MSVC contain no symbol table at all; only exports and imports are named in the file. (PE files built with MinGW are different — they often carry a COFF symbol table and even DWARF unless stripped, which is a gift when you meet one.)

Tip: Always check whether a sample kept its symbols before diving in. On Linux, file reports "not stripped"; on Windows, look for a debug directory, and for MinGW binaries a non-zero symbol count in the COFF file header.

PDB paths: the author's fingerprint

When MSVC links with /DEBUG, it writes a small debug directory entry into the PE — a CodeView "RSDS" record holding a GUID, an age counter and the full path of the PDB file on the build machine. The PDB itself stays behind, but the path ships with the binary:

text
C:\Users\dev\source\repos\Updater\x64\Release\Updater.pdb

These strings leak user names, project names and folder structures. Threat intelligence teams cluster samples by PDB path, and YARA rules often match on them. They can be forged or removed, so treat them as a strong hint rather than proof.

Other build artefacts play a similar role. The undocumented Rich header that Microsoft's linker places after the DOS stub records which compiler and linker versions built each object file, and its hash is widely used to link samples built on the same toolchain.

Why this matters for analysts

  • Library code is noise to filter, not code to read. Recognising that a function is memcpy or part of the C runtime saves hours. Signature systems do this automatically: IDA's FLIRT and Ghidra's Function ID hash known library functions and rename matches in your sample.
  • Missing names are expected, not suspicious. A stripped ELF or an MSVC PE without a PDB is normal. What is suspicious is a binary with almost no imports, which suggests imports are being resolved by hand — see dynamic import resolution.
  • Compiler choices change the code shape. Optimisation level, compiler vendor and language all leave recognisable idioms, which a later module covers in detail.
  • Build metadata is intelligence. PDB paths, Rich headers and compiler version strings connect samples to each other and sometimes to their authors.

Lab: stop the build at every stage

You need a Linux, WSL or container shell with gcc and binutils. An optional Windows section uses the Visual Studio Developer Command Prompt.

  1. Create a small benign program with an external call, a global and a local string:

    c
    // beacon.c — harmless: it only prints
    #include <stdio.h>
    #define HOST "telemetry.example.net"
    
    int interval = 60;
    
    static void report(const char *host) {
        printf("would contact %s every %d s\n", host, interval);
    }
    
    int main(void) {
        report(HOST);
        return 0;
    }
  2. Preprocess and look for your macro:

    bash
    gcc -E beacon.c -o beacon.i
    wc -l beacon.i
    grep -n "telemetry" beacon.i

    HOST is gone; only the literal string remains, and the header has swollen the file to hundreds of lines.

  3. Compile to assembly in Intel syntax:

    bash
    gcc -O0 -S -masm=intel beacon.i -o beacon.s

    Find the labels for main, report and the string literal (usually .LC0). Recompile with -O2 and compare: report has most likely been inlined into main and no longer exists on its own. (On distributions that enable _FORTIFY_SOURCE by default, such as Ubuntu, you may also see printf replaced by __printf_chk at -O2 — another way the compiler rewrites what you wrote.)

  4. Assemble and inspect the object file's holes:

    bash
    gcc -O0 -c beacon.c -o beacon.o
    objdump -d -r -M intel beacon.o
    readelf -r beacon.o
    nm beacon.o

    In the objdump output, the call to printf is followed by zeros and an R_X86_64_PLT32 printf-0x4 line. In nm, main is T (defined in text), interval is D (initialised data), report is t (lowercase: local, because it is static), and printf is U — undefined.

  5. Link and inspect again:

    bash
    gcc beacon.o -o beacon
    objdump -d -M intel beacon | grep -A12 "<main>:"
    nm beacon | grep -E " (main|interval|report)$"
    nm -D beacon

    The holes are filled: the call now targets a real address (a PLT stub), and every symbol has a final address. nm -D lists the dynamic symbols, including printf.

  6. Strip and see what survives:

    bash
    cp beacon beacon.stripped && strip beacon.stripped
    nm beacon.stripped
    nm -D beacon.stripped
    file beacon beacon.stripped

    nm reports no symbols, but nm -D still shows printf — the dynamic loader needs it.

  7. Compare static linking:

    bash
    gcc -static -O0 beacon.c -o beacon.static && strip beacon.static
    ls -l beacon.stripped beacon.static
    nm -D beacon.static

    The static binary is many times larger and has no dynamic symbols at all.

  8. (Optional, Windows) In a Developer Command Prompt:

    bat
    cl /P beacon.c
    cl /c /FA /Od beacon.c
    dumpbin /relocations /symbols beacon.obj
    cl /Zi /Od beacon.c /link /DEBUG
    dumpbin /headers beacon.exe | findstr /i "pdb"

    The last command shows the debug directory entry and the full PDB path baked into the executable.

Questions to answer: Which of main, report, interval and printf could you still name in the stripped dynamic binary, and why? In the stripped static binary, how would you find the code that prints the message? What does your own PDB path reveal about your machine?

Key takeaways

  • A build is preprocess, compile, assemble, link; each stage can be stopped and inspected with gcc -E/-S/-c or cl /P /FA /c.
  • Macros and disabled #ifdef features leave no trace; optimisation and inlining reshape functions beyond recognition.
  • Object files contain code with holes, a symbol table and relocations; the linker resolves symbols, lays out sections and fills the holes.
  • Stripping removes symbols and debug info but not dynamic imports and exports; MSVC PEs keep symbols in a separate PDB, yet leak its path.
  • Static linking buries the author's code in library code; FLIRT and Function ID exist to filter that noise.