Skip to content

Leçon 2.2 · Formats binaires· 40 min

PE Sections and Memory Layout

How the section table maps a PE file into memory, how to convert RVAs to file offsets, and which section anomalies betray packers and loaders.

Cette leçon n’est disponible qu’en anglais pour le moment.

Objectifs

  • Read every field of IMAGE_SECTION_HEADER and decode section permissions
  • Convert between RVAs, virtual addresses and file offsets by hand
  • Explain the purpose of the common sections produced by MSVC and mingw-w64
  • Flag packer indicators: raw/virtual size mismatches, RWX sections, odd names and high entropy

The headers from PE Headers describe the file as a whole. The section table describes its parts: which bytes are code, which are read-only data, which are writable globals, and — crucially — where each part lives on disk versus in memory. Those two layouts are different, and almost every confusion beginners have with PE files comes from mixing them up.

For malware analysis, sections are also where packers leave their fingerprints. A section that is empty on disk but huge in memory, writable and executable at the same time, and filled with random-looking bytes is the textbook shape of a packed sample.

Two layouts, one file

On disk, sections are packed tightly, aligned to FileAlignment (usually 0x200, 512 bytes). In memory, the loader spreads them out on SectionAlignment boundaries (usually 0x1000, one 4 KB page) so that each section can get its own page protections.

text
        ON DISK (file offsets)              IN MEMORY (RVAs from ImageBase)
  0x0000 ┌──────────────────┐          0x0000 ┌──────────────────┐
         │ headers          │                 │ headers          │ R--
  0x0400 ├──────────────────┤          0x1000 ├──────────────────┤
         │ .text   (0x1C00) │ ───────────────►│ .text   (0x1A20) │ R-X
  0x2000 ├──────────────────┤                 │   ...padding     │
         │ .data   (0x200)  │ ──┐      0x3000 ├──────────────────┤
  0x2200 ├──────────────────┤   └────────────►│ .data   (0xA0)   │ RW-
         │ .rdata  (0xA00)  │ ──┐      0x4000 ├──────────────────┤
  0x2C00 ├──────────────────┤   └────────────►│ .rdata  (0x9C8)  │ R--
         │ ...              │                 │ ...              │
         └──────────────────┘                 └──────────────────┘

The numbers above come from a 64-bit mingw-w64 build of a hello-world program — the same binary you will build in the lab. Notice that .text starts at file offset 0x400 but at RVA 0x1000. A byte's position in the file tells you nothing about its address in memory until you consult the section table.

Tip: When a memory dump of a process looks "misaligned" in a PE parser, it is usually because the dump is in memory layout while the tool expects file layout. Tools such as PE-bear and pe-sieve can remap ("unmap") a dump back to file layout.

The section header

The section table starts right after the optional header, at e_lfanew + 4 + 20 + SizeOfOptionalHeader, and contains NumberOfSections entries of 40 bytes each:

OffsetFieldSizeMeaning
0x00Name8 bytesASCII, NUL-padded, not necessarily NUL-terminated
0x08VirtualSizeDWORDSize in memory (unaligned)
0x0CVirtualAddressDWORDRVA where the section starts in memory
0x10SizeOfRawDataDWORDSize on disk, a multiple of FileAlignment
0x14PointerToRawDataDWORDFile offset of the section data
0x18PointerToRelocationsDWORDObject files only; 0 in images
0x1CPointerToLinenumbersDWORDDeprecated; 0
0x20NumberOfRelocationsWORDObject files only
0x22NumberOfLinenumbersWORDDeprecated
0x24CharacteristicsDWORDContent type and memory permissions

(VirtualSize is formally Misc.VirtualSize in winnt.h, which is why pefile calls it Misc_VirtualSize.)

Two rules govern how the loader uses these fields:

  • It copies SizeOfRawData bytes from PointerToRawData to ImageBase + VirtualAddress.
  • If VirtualSize is larger than SizeOfRawData, the rest of the section is zero-filled. That is how uninitialised data (.bss) costs nothing on disk.

Because SizeOfRawData is rounded up to FileAlignment and VirtualSize is not, small differences in either direction are normal. .text above has SizeOfRawData = 0x1C00 but VirtualSize = 0x1A20: the extra 0x1E0 bytes on disk are padding.

Characteristics: what and how

The high bits are the memory permissions the loader applies; the low bits describe content.

FlagValueMeaning
CNT_CODE0x00000020Contains code
CNT_INITIALIZED_DATA0x00000040Contains initialised data
CNT_UNINITIALIZED_DATA0x00000080Contains uninitialised data
MEM_DISCARDABLE0x02000000Can be discarded after load (e.g. .reloc)
MEM_SHARED0x10000000Shared between processes
MEM_EXECUTE0x20000000Executable
MEM_READ0x40000000Readable
MEM_WRITE0x80000000Writable

So 0x60000020 is code, read + execute; 0xC0000040 is initialised data, read + write; 0x40000040 is read-only data. Any value with both 0x20000000 and 0x80000000 set — for example 0xE0000060 or 0xE0000080 — is RWX. See memory protection for how these map to page protections.

Common sections

Section names are pure convention. The loader never looks at them; it uses the data directories to find imports, resources and so on. Still, compilers are consistent, so the names are a quick orientation guide.

NameTypical permissionsContents
.textR-XCompiled code, including CRT startup code
.rdataR--Constants, string literals, often the import and export tables (MSVC)
.dataRW-Initialised global and static variables
.bssRW-Uninitialised globals; SizeOfRawData is 0 (mingw)
.idataR-- or RW-Import tables in their own section (mingw, older linkers)
.edataR--Export table in its own section (mingw DLLs)
.pdataR--x64 exception/unwind function table (RUNTIME_FUNCTION)
.xdataR--x64 unwind info (mingw)
.rsrcR--Resource tree: icons, manifests, version info, embedded files
.relocR--, discardableBase relocations for loading at a non-preferred base
.tlsRW-Thread-local storage template
.CRTR--C runtime initialiser tables (MSVC)

Deviations tell stories. UPX0/UPX1, .aspack, .themida, .vmp0, .enigma1 and .MPRESS1 are packer or protector names. Random strings, empty names or names like .text appearing twice suggest a custom packer or a hand-built file. Because the loader ignores names, though, a careful author can call their encrypted payload .rdata — so names are hints, never proof.

RVA ↔ file offset: a worked example

You will do this conversion constantly: a debugger shows you an address, and you want to patch or carve the bytes in the file; or a header gives you an RVA and you want to find it with xxd.

The formula, for the section that contains the RVA:

text
file_offset = RVA - section.VirtualAddress + section.PointerToRawData
VA          = ImageBase + RVA

A section contains an RVA if VirtualAddress <= RVA < VirtualAddress + VirtualSize. Take our hello-world build, with ImageBase = 0x140000000:

SectionVirtualAddressVirtualSizePointerToRawDataSizeOfRawData
.text0x10000x1A200x4000x1C00
.data0x30000xA00x20000x200
.rdata0x40000x9C80x22000xA00
.bss0x70000x1800x00x0

Example 1 — the entry point. AddressOfEntryPoint = 0x1440. It falls in .text (0x1000 ≤ 0x1440 < 0x2A20):

text
file_offset = 0x1440 - 0x1000 + 0x400 = 0x840
VA          = 0x140000000 + 0x1440 = 0x140001440

xxd -s 0x840 -l 16 hello.exe shows the first bytes of the startup code, and a debugger will break at 0x140001440 if ASLR does not move the image.

Example 2 — the TLS directory. Data directory 9 says RVA 0x4040. That is in .rdata: 0x4040 - 0x4000 + 0x2200 = 0x2240.

Example 3 — a trap. RVA 0x7010 is in .bss. It has no file offset at all: those bytes exist only in memory, as zeros. The same applies to the part of any section beyond SizeOfRawData. Code that the unpacker writes into such a region at runtime cannot be found in the file — you must dump memory.

Warning: Real parsers must also handle headers (RVAs below the first section map 1:1 to file offsets), PointerToRawData values that are not aligned (the loader rounds them down to 512), and sections that overlap. Malware sometimes relies on exactly these edge cases to confuse tools. In pefile, use pe.get_offset_from_rva() rather than writing your own for anything serious.

Section anomalies and what they suggest

Virtual size much larger than raw size

A packer compresses the original program into one section and reserves an empty section large enough to hold it once decompressed. The classic example is UPX: UPX0 has SizeOfRawData = 0 but a VirtualSize equal to the unpacked image, and UPX1 contains the compressed data plus the decompression stub. At runtime, the stub decompresses into UPX0, fixes imports and jumps to the original entry point (OEP). See UPX packing and runtime crypters for the wider family.

A .bss-style section with raw size 0 is normal. An executable section with raw size 0 is not: code that is not in the file must be written there at runtime.

Writable and executable sections

Normal compilers never emit RWX sections. Code is R-X, data is RW-. A section with both write and execute permissions exists because something intends to write code and then run it — an unpacker, a crypter or self-modifying code. It is one of the strongest static packer indicators.

Entry point in an unexpected section

The entry point normally lands in .text (or the first executable section). An entry point in the last section, in .rsrc, in a section with a packer name, or in a writable section points at an unpacking stub. See entry point obfuscation.

High entropy

Shannon entropy measures how unpredictable the bytes are, from 0 (all bytes identical) to 8 bits per byte (perfectly random). Rough guidance:

EntropyTypical content
0 – 1Zeros, padding, empty sections
4.5 – 6.5Native x86/x64 code, text, tables
6.5 – 7.2Dense code, some resources
> 7.2Compressed or encrypted data

Always compute entropy per section, not for the whole file: a large, uncompressed .text can hide a small encrypted blob in .data. Also remember that images, compressed installers and certificates are legitimately high-entropy. Detect It Easy and PE-bear both draw entropy graphs; the entropy glossary entry covers the maths.

Other oddities

  • Section data extending past the end of the file (truncated or corrupted sample).
  • Gaps between sections on disk, or data after the last section — an overlay, covered in Resources, Overlays and Other Hiding Places.
  • A SizeOfImage smaller than the end of the last section.
  • A TLS section or directory in a small program that has no reason to use thread-local storage — check for TLS callbacks.

Lab: a section triage script

You will write a small section analyser, run it on a benign program, then pack that program with UPX and watch the indicators appear. UPX is a legitimate, open-source packer and the result is still your harmless program.

  1. Build the hello-world binary from the previous lesson if you do not have it:

    bash
    x86_64-w64-mingw32-gcc -O0 -s -o hello.exe hello.c
  2. Install the tools: pip install pefile, and UPX from your package manager (apt install upx-ucl or brew install upx).

  3. Save the script:

    python
    # sections.py
    import math, sys
    import pefile
    
    STANDARD = {".text", ".rdata", ".data", ".pdata", ".xdata", ".rsrc",
                ".reloc", ".tls", ".bss", ".idata", ".edata", ".CRT"}
    R, W, X = 0x40000000, 0x80000000, 0x20000000
    
    def entropy(buf: bytes) -> float:
        if not buf:
            return 0.0
        counts = [0] * 256
        for b in buf:
            counts[b] += 1
        n = len(buf)
        return abs(-sum(c / n * math.log2(c / n) for c in counts if c))
    
    pe = pefile.PE(sys.argv[1])
    ep = pe.OPTIONAL_HEADER.AddressOfEntryPoint
    print(f"{'name':8} {'VA':>8} {'VSize':>8} {'RawOff':>8} {'RawSize':>8} perm  {'H':>5}  flags")
    for s in pe.sections:
        name = s.Name.rstrip(b"\x00").decode(errors="replace")
        c = s.Characteristics
        perm = ("R" if c & R else "-") + ("W" if c & W else "-") + ("X" if c & X else "-")
        h = entropy(s.get_data())
        flags = []
        if c & W and c & X:
            flags.append("RWX")
        if s.SizeOfRawData == 0 and s.Misc_VirtualSize > 0 and c & X:
            flags.append("exec-but-empty-on-disk")
        elif s.SizeOfRawData and s.Misc_VirtualSize > 4 * s.SizeOfRawData:
            flags.append("vsize>>raw")
        if name not in STANDARD:
            flags.append("odd-name")
        if h > 7.2:
            flags.append("high-entropy")
        if s.contains_rva(ep):
            flags.append("<-EP")
        print(f"{name:8} {s.VirtualAddress:#8x} {s.Misc_VirtualSize:#8x} "
              f"{s.PointerToRawData:#8x} {s.SizeOfRawData:#8x} {perm}   {h:5.2f}  "
              + " ".join(flags))
  4. Run it on the unpacked binary:

    bash
    python3 sections.py hello.exe
    text
    name           VA    VSize   RawOff  RawSize perm      H  flags
    .text      0x1000   0x1a20    0x400   0x1c00 R-X    5.69  <-EP
    .data      0x3000     0xa0   0x2000    0x200 RW-    0.57
    .rdata     0x4000    0x9c8   0x2200    0xa00 R--    4.03
    .pdata     0x5000    0x21c   0x2c00    0x400 R--    2.30
    .xdata     0x6000    0x1a4   0x3000    0x200 R--    3.31
    .bss       0x7000    0x180      0x0      0x0 RW-    0.00
    .idata     0x8000    0x874   0x3200    0xa00 R--    3.52
    .tls       0x9000     0x10   0x3c00    0x200 RW-    0.00
    .reloc     0xa000     0x60   0x3e00    0x200 R--    1.24

    Exact numbers depend on your mingw-w64 version. The shape is what matters: standard names, no RWX, entry point in .text, entropy below 6.

  5. Pack a copy and run the script again:

    bash
    upx -o hello_upx.exe hello.exe
    python3 sections.py hello_upx.exe
    text
    name           VA    VSize   RawOff  RawSize perm      H  flags
    UPX0       0x1000   0xa000    0x200      0x0 RWX    0.00  RWX exec-but-empty-on-disk odd-name
    UPX1       0xb000   0x2000    0x200   0x1c00 RWX    7.49  RWX odd-name high-entropy <-EP
    UPX2       0xd000   0x1000   0x1e00    0x400 RW-    3.31  odd-name

    Depending on the UPX version and whether the input has resources, the third section may be named .rsrc instead of UPX2.

  6. Convert an RVA by hand. Take the entry point RVA of each file from pefile (pe.OPTIONAL_HEADER.AddressOfEntryPoint), compute its file offset with the formula above, then check your answer with pe.get_offset_from_rva(rva) and look at the bytes with xxd -s.

  7. Look at both files in PE-bear. Open the Section Hdrs tab and compare the raw and virtual columns, then look at the section map. Note how UPX0 occupies address space but no file space.

  8. Unpack and compare: upx -d -o hello_unpacked.exe hello_upx.exe, then run the script once more.

Questions to answer: Why does UPX0 need to be writable and executable? What is the SizeOfImage of the packed file compared with the original, and why is it at least as large? Which single indicator in your output would you trust least on its own, and why? Is hello_unpacked.exe byte-identical to hello.exe (compare the hashes)?

Key takeaways

  • Each 40-byte section header maps SizeOfRawData bytes at PointerToRawData on disk to VirtualAddress in memory; anything beyond raw size up to VirtualSize is zero-filled.
  • file_offset = RVA − VirtualAddress + PointerToRawData, for the section containing the RVA; RVAs in zero-filled regions have no file offset.
  • Section names are convention only; permissions come from Characteristics.
  • Packer indicators: executable sections empty on disk, VirtualSize far above raw size, RWX permissions, unusual names, entry point outside .text and per-section entropy above about 7.2.
  • No single indicator is proof — combine them, then confirm dynamically.