Leçon 2.2 · Formats binaires· 40 min
PE Sections and Memory Layout
How the section table maps a PE file into memory, how to convert RVAs to file offsets, and which section anomalies betray packers and loaders.
Cette leçon n’est disponible qu’en anglais pour le moment.
Objectifs
- Read every field of IMAGE_SECTION_HEADER and decode section permissions
- Convert between RVAs, virtual addresses and file offsets by hand
- Explain the purpose of the common sections produced by MSVC and mingw-w64
- Flag packer indicators: raw/virtual size mismatches, RWX sections, odd names and high entropy
The headers from PE Headers describe the file as a whole. The section table describes its parts: which bytes are code, which are read-only data, which are writable globals, and — crucially — where each part lives on disk versus in memory. Those two layouts are different, and almost every confusion beginners have with PE files comes from mixing them up.
For malware analysis, sections are also where packers leave their fingerprints. A section that is empty on disk but huge in memory, writable and executable at the same time, and filled with random-looking bytes is the textbook shape of a packed sample.
Two layouts, one file
On disk, sections are packed tightly, aligned to FileAlignment (usually
0x200, 512 bytes). In memory, the loader spreads them out on
SectionAlignment boundaries (usually 0x1000, one 4 KB page) so that each
section can get its own page protections.
ON DISK (file offsets) IN MEMORY (RVAs from ImageBase)
0x0000 ┌──────────────────┐ 0x0000 ┌──────────────────┐
│ headers │ │ headers │ R--
0x0400 ├──────────────────┤ 0x1000 ├──────────────────┤
│ .text (0x1C00) │ ───────────────►│ .text (0x1A20) │ R-X
0x2000 ├──────────────────┤ │ ...padding │
│ .data (0x200) │ ──┐ 0x3000 ├──────────────────┤
0x2200 ├──────────────────┤ └────────────►│ .data (0xA0) │ RW-
│ .rdata (0xA00) │ ──┐ 0x4000 ├──────────────────┤
0x2C00 ├──────────────────┤ └────────────►│ .rdata (0x9C8) │ R--
│ ... │ │ ... │
└──────────────────┘ └──────────────────┘The numbers above come from a 64-bit mingw-w64 build of a hello-world program —
the same binary you will build in the lab. Notice that .text starts at file
offset 0x400 but at RVA 0x1000. A byte's position in the file tells you
nothing about its address in memory until you consult the section table.
Tip: When a memory dump of a process looks "misaligned" in a PE parser, it is usually because the dump is in memory layout while the tool expects file layout. Tools such as PE-bear and pe-sieve can remap ("unmap") a dump back to file layout.
The section header
The section table starts right after the optional header, at
e_lfanew + 4 + 20 + SizeOfOptionalHeader, and contains NumberOfSections
entries of 40 bytes each:
| Offset | Field | Size | Meaning |
|---|---|---|---|
0x00 | Name | 8 bytes | ASCII, NUL-padded, not necessarily NUL-terminated |
0x08 | VirtualSize | DWORD | Size in memory (unaligned) |
0x0C | VirtualAddress | DWORD | RVA where the section starts in memory |
0x10 | SizeOfRawData | DWORD | Size on disk, a multiple of FileAlignment |
0x14 | PointerToRawData | DWORD | File offset of the section data |
0x18 | PointerToRelocations | DWORD | Object files only; 0 in images |
0x1C | PointerToLinenumbers | DWORD | Deprecated; 0 |
0x20 | NumberOfRelocations | WORD | Object files only |
0x22 | NumberOfLinenumbers | WORD | Deprecated |
0x24 | Characteristics | DWORD | Content type and memory permissions |
(VirtualSize is formally Misc.VirtualSize in winnt.h, which is why pefile
calls it Misc_VirtualSize.)
Two rules govern how the loader uses these fields:
- It copies
SizeOfRawDatabytes fromPointerToRawDatatoImageBase + VirtualAddress. - If
VirtualSizeis larger thanSizeOfRawData, the rest of the section is zero-filled. That is how uninitialised data (.bss) costs nothing on disk.
Because SizeOfRawData is rounded up to FileAlignment and VirtualSize is
not, small differences in either direction are normal. .text above has
SizeOfRawData = 0x1C00 but VirtualSize = 0x1A20: the extra 0x1E0 bytes on
disk are padding.
Characteristics: what and how
The high bits are the memory permissions the loader applies; the low bits describe content.
| Flag | Value | Meaning |
|---|---|---|
CNT_CODE | 0x00000020 | Contains code |
CNT_INITIALIZED_DATA | 0x00000040 | Contains initialised data |
CNT_UNINITIALIZED_DATA | 0x00000080 | Contains uninitialised data |
MEM_DISCARDABLE | 0x02000000 | Can be discarded after load (e.g. .reloc) |
MEM_SHARED | 0x10000000 | Shared between processes |
MEM_EXECUTE | 0x20000000 | Executable |
MEM_READ | 0x40000000 | Readable |
MEM_WRITE | 0x80000000 | Writable |
So 0x60000020 is code, read + execute; 0xC0000040 is initialised data, read +
write; 0x40000040 is read-only data. Any value with both 0x20000000 and
0x80000000 set — for example 0xE0000060 or 0xE0000080 — is RWX. See
memory protection for how these map to page
protections.
Common sections
Section names are pure convention. The loader never looks at them; it uses the data directories to find imports, resources and so on. Still, compilers are consistent, so the names are a quick orientation guide.
| Name | Typical permissions | Contents |
|---|---|---|
.text | R-X | Compiled code, including CRT startup code |
.rdata | R-- | Constants, string literals, often the import and export tables (MSVC) |
.data | RW- | Initialised global and static variables |
.bss | RW- | Uninitialised globals; SizeOfRawData is 0 (mingw) |
.idata | R-- or RW- | Import tables in their own section (mingw, older linkers) |
.edata | R-- | Export table in its own section (mingw DLLs) |
.pdata | R-- | x64 exception/unwind function table (RUNTIME_FUNCTION) |
.xdata | R-- | x64 unwind info (mingw) |
.rsrc | R-- | Resource tree: icons, manifests, version info, embedded files |
.reloc | R--, discardable | Base relocations for loading at a non-preferred base |
.tls | RW- | Thread-local storage template |
.CRT | R-- | C runtime initialiser tables (MSVC) |
Deviations tell stories. UPX0/UPX1, .aspack, .themida, .vmp0, .enigma1
and .MPRESS1 are packer or protector names. Random strings, empty names or
names like .text appearing twice suggest a custom packer or a hand-built file.
Because the loader ignores names, though, a careful author can call their
encrypted payload .rdata — so names are hints, never proof.
RVA ↔ file offset: a worked example
You will do this conversion constantly: a debugger shows you an address, and you
want to patch or carve the bytes in the file; or a header gives you an RVA and
you want to find it with xxd.
The formula, for the section that contains the RVA:
file_offset = RVA - section.VirtualAddress + section.PointerToRawData
VA = ImageBase + RVAA section contains an RVA if
VirtualAddress <= RVA < VirtualAddress + VirtualSize. Take our hello-world
build, with ImageBase = 0x140000000:
| Section | VirtualAddress | VirtualSize | PointerToRawData | SizeOfRawData |
|---|---|---|---|---|
.text | 0x1000 | 0x1A20 | 0x400 | 0x1C00 |
.data | 0x3000 | 0xA0 | 0x2000 | 0x200 |
.rdata | 0x4000 | 0x9C8 | 0x2200 | 0xA00 |
.bss | 0x7000 | 0x180 | 0x0 | 0x0 |
Example 1 — the entry point. AddressOfEntryPoint = 0x1440. It falls in
.text (0x1000 ≤ 0x1440 < 0x2A20):
file_offset = 0x1440 - 0x1000 + 0x400 = 0x840
VA = 0x140000000 + 0x1440 = 0x140001440xxd -s 0x840 -l 16 hello.exe shows the first bytes of the startup code, and a
debugger will break at 0x140001440 if ASLR does not move the image.
Example 2 — the TLS directory. Data directory 9 says RVA 0x4040. That is
in .rdata: 0x4040 - 0x4000 + 0x2200 = 0x2240.
Example 3 — a trap. RVA 0x7010 is in .bss. It has no file offset at all:
those bytes exist only in memory, as zeros. The same applies to the part of any
section beyond SizeOfRawData. Code that the unpacker writes into such a region
at runtime cannot be found in the file — you must dump memory.
Warning: Real parsers must also handle headers (RVAs below the first section map 1:1 to file offsets),
PointerToRawDatavalues that are not aligned (the loader rounds them down to 512), and sections that overlap. Malware sometimes relies on exactly these edge cases to confuse tools. In pefile, usepe.get_offset_from_rva()rather than writing your own for anything serious.
Section anomalies and what they suggest
Virtual size much larger than raw size
A packer compresses the original program into one section and reserves an
empty section large enough to hold it once decompressed. The classic example is
UPX: UPX0 has SizeOfRawData = 0 but a VirtualSize equal to the unpacked
image, and UPX1 contains the compressed data plus the decompression stub. At
runtime, the stub decompresses into UPX0, fixes imports and jumps to the
original entry point (OEP). See
UPX packing and
runtime crypters for the wider family.
A .bss-style section with raw size 0 is normal. An executable section with
raw size 0 is not: code that is not in the file must be written there at
runtime.
Writable and executable sections
Normal compilers never emit RWX sections. Code is R-X, data is RW-. A section with both write and execute permissions exists because something intends to write code and then run it — an unpacker, a crypter or self-modifying code. It is one of the strongest static packer indicators.
Entry point in an unexpected section
The entry point normally lands in .text (or the first executable section).
An entry point in the last section, in .rsrc, in a section with a packer name,
or in a writable section points at an unpacking stub. See
entry point obfuscation.
High entropy
Shannon entropy measures how unpredictable the bytes are, from 0 (all bytes identical) to 8 bits per byte (perfectly random). Rough guidance:
| Entropy | Typical content |
|---|---|
| 0 – 1 | Zeros, padding, empty sections |
| 4.5 – 6.5 | Native x86/x64 code, text, tables |
| 6.5 – 7.2 | Dense code, some resources |
| > 7.2 | Compressed or encrypted data |
Always compute entropy per section, not for the whole file: a large,
uncompressed .text can hide a small encrypted blob in .data. Also remember
that images, compressed installers and certificates are legitimately
high-entropy. Detect It Easy and PE-bear both draw entropy graphs; the
entropy glossary entry covers the maths.
Other oddities
- Section data extending past the end of the file (truncated or corrupted sample).
- Gaps between sections on disk, or data after the last section — an overlay, covered in Resources, Overlays and Other Hiding Places.
- A
SizeOfImagesmaller than the end of the last section. - A TLS section or directory in a small program that has no reason to use thread-local storage — check for TLS callbacks.
Lab: a section triage script
You will write a small section analyser, run it on a benign program, then pack that program with UPX and watch the indicators appear. UPX is a legitimate, open-source packer and the result is still your harmless program.
-
Build the hello-world binary from the previous lesson if you do not have it:
bash x86_64-w64-mingw32-gcc -O0 -s -o hello.exe hello.c -
Install the tools:
pip install pefile, and UPX from your package manager (apt install upx-uclorbrew install upx). -
Save the script:
python # sections.py import math, sys import pefile STANDARD = {".text", ".rdata", ".data", ".pdata", ".xdata", ".rsrc", ".reloc", ".tls", ".bss", ".idata", ".edata", ".CRT"} R, W, X = 0x40000000, 0x80000000, 0x20000000 def entropy(buf: bytes) -> float: if not buf: return 0.0 counts = [0] * 256 for b in buf: counts[b] += 1 n = len(buf) return abs(-sum(c / n * math.log2(c / n) for c in counts if c)) pe = pefile.PE(sys.argv[1]) ep = pe.OPTIONAL_HEADER.AddressOfEntryPoint print(f"{'name':8} {'VA':>8} {'VSize':>8} {'RawOff':>8} {'RawSize':>8} perm {'H':>5} flags") for s in pe.sections: name = s.Name.rstrip(b"\x00").decode(errors="replace") c = s.Characteristics perm = ("R" if c & R else "-") + ("W" if c & W else "-") + ("X" if c & X else "-") h = entropy(s.get_data()) flags = [] if c & W and c & X: flags.append("RWX") if s.SizeOfRawData == 0 and s.Misc_VirtualSize > 0 and c & X: flags.append("exec-but-empty-on-disk") elif s.SizeOfRawData and s.Misc_VirtualSize > 4 * s.SizeOfRawData: flags.append("vsize>>raw") if name not in STANDARD: flags.append("odd-name") if h > 7.2: flags.append("high-entropy") if s.contains_rva(ep): flags.append("<-EP") print(f"{name:8} {s.VirtualAddress:#8x} {s.Misc_VirtualSize:#8x} " f"{s.PointerToRawData:#8x} {s.SizeOfRawData:#8x} {perm} {h:5.2f} " + " ".join(flags)) -
Run it on the unpacked binary:
bash python3 sections.py hello.exetext name VA VSize RawOff RawSize perm H flags .text 0x1000 0x1a20 0x400 0x1c00 R-X 5.69 <-EP .data 0x3000 0xa0 0x2000 0x200 RW- 0.57 .rdata 0x4000 0x9c8 0x2200 0xa00 R-- 4.03 .pdata 0x5000 0x21c 0x2c00 0x400 R-- 2.30 .xdata 0x6000 0x1a4 0x3000 0x200 R-- 3.31 .bss 0x7000 0x180 0x0 0x0 RW- 0.00 .idata 0x8000 0x874 0x3200 0xa00 R-- 3.52 .tls 0x9000 0x10 0x3c00 0x200 RW- 0.00 .reloc 0xa000 0x60 0x3e00 0x200 R-- 1.24Exact numbers depend on your mingw-w64 version. The shape is what matters: standard names, no RWX, entry point in
.text, entropy below 6. -
Pack a copy and run the script again:
bash upx -o hello_upx.exe hello.exe python3 sections.py hello_upx.exetext name VA VSize RawOff RawSize perm H flags UPX0 0x1000 0xa000 0x200 0x0 RWX 0.00 RWX exec-but-empty-on-disk odd-name UPX1 0xb000 0x2000 0x200 0x1c00 RWX 7.49 RWX odd-name high-entropy <-EP UPX2 0xd000 0x1000 0x1e00 0x400 RW- 3.31 odd-nameDepending on the UPX version and whether the input has resources, the third section may be named
.rsrcinstead ofUPX2. -
Convert an RVA by hand. Take the entry point RVA of each file from
pefile(pe.OPTIONAL_HEADER.AddressOfEntryPoint), compute its file offset with the formula above, then check your answer withpe.get_offset_from_rva(rva)and look at the bytes withxxd -s. -
Look at both files in PE-bear. Open the Section Hdrs tab and compare the raw and virtual columns, then look at the section map. Note how
UPX0occupies address space but no file space. -
Unpack and compare:
upx -d -o hello_unpacked.exe hello_upx.exe, then run the script once more.
Questions to answer: Why does UPX0 need to be writable and executable?
What is the SizeOfImage of the packed file compared with the original, and why
is it at least as large? Which single indicator in your output would you trust
least on its own, and why? Is hello_unpacked.exe byte-identical to
hello.exe (compare the hashes)?
Key takeaways
- Each 40-byte section header maps
SizeOfRawDatabytes atPointerToRawDataon disk toVirtualAddressin memory; anything beyond raw size up toVirtualSizeis zero-filled. file_offset = RVA − VirtualAddress + PointerToRawData, for the section containing the RVA; RVAs in zero-filled regions have no file offset.- Section names are convention only; permissions come from
Characteristics. - Packer indicators: executable sections empty on disk,
VirtualSizefar above raw size, RWX permissions, unusual names, entry point outside.textand per-section entropy above about 7.2. - No single indicator is proof — combine them, then confirm dynamically.