Leçon 3.4 · Triage statique· 35 min
Detecting Packers and Entropy
Recognise packed and encrypted executables with entropy, section anomalies, import tables and signatures — and avoid the classic false positives.
Cette leçon n’est disponible qu’en anglais pour le moment.
Objectifs
- Explain what packers, crypters and protectors do and why they defeat static analysis
- Compute Shannon entropy and interpret whole-file, per-section and sliding-window values
- Combine entropy with structural indicators to reach a confident packing verdict
- Recognise legitimate sources of high entropy that are not packing
Everything in this module so far — strings, imports, capabilities — assumes the interesting code is sitting in the file where you can read it. A packed sample breaks that assumption. Its real code is compressed or encrypted, and a small stub restores it in memory at runtime. Static triage of a packed file mostly describes the packer, not the malware.
So one of the first questions of any triage is: is this packed? This lesson teaches you to answer it with evidence, and to avoid the false alarms that trap beginners.
Packers, crypters and protectors
The terms overlap, but they point at different goals:
| Kind | Main goal | Examples | Typical stub behaviour |
|---|---|---|---|
| Packer (compressor) | Smaller files | UPX, MPRESS, ASPack | Decompress sections, rebuild imports, jump to the original entry point |
| Crypter | Evade signature detection | Commodity crypters sold with malware | Decrypt an embedded payload, often into a new process or fresh memory |
| Protector | Resist reversing and tampering | Themida, VMProtect, Enigma | Anti-debugging, import obfuscation, code virtualization |
All three share a runtime pattern:
on disk in memory at runtime
┌───────────────┐ ┌────────────────────┐
│ headers │ │ headers │
│ stub code │ ── stub runs ───► │ stub code │
│ packed blob │ decompresses / │ original code ◄───┼── jump to OEP
│ (high entropy)│ decrypts, │ original data │
└───────────────┘ fixes imports │ rebuilt IAT │
└────────────────────┘The stub ends by transferring control to the original entry point (OEP) of the unpacked code. Finding that moment and dumping memory is the job of the upcoming Unpacking module; this lesson is only about detection. The UPX packing and runtime crypter technique pages describe both families in more depth.
Entropy in one formula
Shannon entropy measures how unpredictable a block of bytes is. For a buffer where byte value i appears with probability pᵢ:
H = − Σ pᵢ · log₂(pᵢ) summed over the 256 possible byte valuesThe result is in bits per byte, from 0.0 to 8.0:
- 0.0 — every byte is identical (a block of zeros).
- ≈ 3–5 — text, tables, structured data.
- ≈ 5–6.5 — typical compiled x86/x64 code.
- ≈ 7.2 and above — compressed or encrypted data; truly random bytes approach 8.0.
Compression and encryption both remove redundancy, which is why they push entropy toward the maximum. Plain code cannot get there: opcodes, registers and common instruction patterns repeat too much.
Tip: Entropy depends on how many bytes you measure. A 1 KiB window of perfectly random data only scores about 7.8, because 1024 samples cannot fill 256 buckets evenly. Compare values measured with the same window size.
Three ways to measure
- Whole file — one number; quick but coarse. A large, low-entropy section can hide a small encrypted blob.
- Per section — the most useful triage view: which part of the file is dense?
- Sliding window — entropy for each consecutive block (say 1–4 KiB), plotted as a graph. Detect It Easy's entropy view and the lab script below both do this. It reveals blobs inside a section and data in the overlay.
Indicators beyond entropy
Entropy alone is never a verdict. Combine it with the structural indicators you learned in PE Sections and Reading Capabilities from Imports:
| Indicator | Why it suggests packing |
|---|---|
| A section with raw size 0 but a large virtual size | Space reserved for code that will be written at runtime (UPX's UPX0) |
| Writable + executable section | The stub must write the code it will later run |
Entry point in the last section or outside .text | Execution starts in the stub, not the program |
Tiny import table with LoadLibrary, GetProcAddress, VirtualProtect / VirtualAlloc | Imports are rebuilt at runtime (see dynamic import resolution) |
Unusual section names (UPX0, .aspack, .themida, random strings) | Packer-generated layout |
| Very few readable strings | Strings are inside the compressed blob |
| High-entropy overlay or resource | Payload stored outside the sections (see Resources and Overlays) |
Three or more of these together, plus a dense section, is a confident verdict.
Signature-based identification
Tools such as Detect It Easy (DiE) match known byte patterns at the entry point and structural rules to name the packer, compiler and linker. That is fast and often exactly right for common packers — but signatures can be spoofed (a sample can include fake UPX section names) and custom crypters have no signature at all. Treat a DiE label as a lead to confirm, and treat "no packer detected" as "no known packer detected". capa likewise reports when a file looks packed and matches rules such as "packed with UPX".
The false positives
High entropy is common in perfectly legitimate files:
- Compressed resources. PNG and JPEG icons, embedded ZIPs and fonts are
already compressed. An
.rsrcsection full of images scores high. - Installers and self-extractors. NSIS, Inno Setup and 7-Zip SFX files carry a compressed archive in the overlay — by design.
- Legitimate protectors. Commercial software, games and licensed tools often use protectors; .NET applications may be obfuscated.
- Embedded certificates and crypto tables — small but dense.
The discriminator is where the entropy is. Dense data in a resource that the program loads as an image is normal. Dense data in a section the entry point lives in, or a blob that is decrypted into executable memory, is not.
Lab: measure, then judge
You need mingw-w64, UPX and Python 3 with pefile. You will compare three
harmless binaries: a plain program, the same program packed with UPX, and a
program that embeds 32 KiB of random bytes as a resource.
-
Build the three variants:
bash printf '#include <stdio.h>\nint main(void){puts("hello, analyst");return 0;}\n' > hello.c x86_64-w64-mingw32-gcc -s -o hello.exe hello.c upx -o hello_upx.exe hello.exe head -c 32768 /dev/urandom > blob.bin echo '101 RCDATA "blob.bin"' > blob.rc x86_64-w64-mingw32-windres blob.rc -O coff -o blob.res.o x86_64-w64-mingw32-gcc -s -o hello_blob.exe hello.c blob.res.o -
Save this script as
entropy.py:python # entropy.py — per-section and sliding-window entropy for a PE file import math, sys from collections import Counter import pefile def entropy(buf: bytes) -> float: """Shannon entropy in bits per byte, 0.0 to 8.0.""" n = len(buf) if n == 0: return 0.0 return max(0.0, -sum(c / n * math.log2(c / n) for c in Counter(buf).values())) def region(pe, off: int) -> str: """Name the part of the file that contains file offset `off`.""" for s in pe.sections: if s.PointerToRawData <= off < s.PointerToRawData + s.SizeOfRawData: return s.Name.rstrip(b"\x00").decode(errors="replace") end = max(s.PointerToRawData + s.SizeOfRawData for s in pe.sections) return "overlay" if off >= end else "headers" path = sys.argv[1] window = int(sys.argv[2]) if len(sys.argv) > 2 else 1024 data = open(path, "rb").read() pe = pefile.PE(data=data) print(f"{path}: {len(data)} bytes, whole-file entropy {entropy(data):.2f}\n") print(f"{'section':8} {'raw size':>9} {'entropy':>8}") for s in pe.sections: name = s.Name.rstrip(b"\x00").decode(errors="replace") print(f"{name:8} {s.SizeOfRawData:>9} {entropy(s.get_data()):>8.2f}") print(f"\nwindow = {window} bytes; one '#' = 0.25 bits/byte; '|' marks 7.2") for off in range(0, len(data), window): h = entropy(data[off:off + window]) bar = list(("#" * round(h * 4)).ljust(32)) bar[29] = "|" print(f"{off:#08x} {region(pe, off):8} {h:4.2f} {''.join(bar)}") -
Run it on the plain program:
bash python3 entropy.py hello.exe 4096text hello.exe: 16384 bytes, whole-file entropy 4.70 section raw size entropy .text 7168 5.69 .data 512 0.57 .rdata 2560 4.02 .pdata 1024 2.30 .xdata 512 3.31 .bss 0 0.00 .idata 2560 3.52 .tls 512 0.00 .reloc 512 1.24Code sits in the expected 5–6.5 band; nothing crosses 7.2.
-
Run it on the UPX build:
text hello_upx.exe: 8704 bytes, whole-file entropy 7.00 section raw size entropy UPX0 0 0.00 UPX1 7168 7.48 UPX2 1024 3.31Note every structural indicator at once:
UPX0has no raw data,UPX1is dense, the section names are a signature, and the import table (runcapimports.pyfrom the previous lesson) is reduced toLoadLibraryA,GetProcAddress,VirtualProtectandExitProcessplus one import per C runtime DLL. In the sliding-window view, the first window is labelledheadersbut scores high: with a 4096-byte window it already spans most ofUPX1. Re-run with1024to see the boundary sharply. -
Run it on the program with the random resource:
text hello_blob.exe: 49664 bytes, whole-file entropy 7.27 ... .text 7168 5.69 .rsrc 33280 7.97The whole-file entropy is higher than the UPX build, yet this program is not packed at all: the code is untouched, the imports are normal, and the dense bytes sit in a resource the program never executes.
-
Optional: open all three in Detect It Easy (entropy graph and signature scan) and in PE-bear, and compare with your script.
Questions to answer: Which single number would have misled you in step 5,
and which indicators corrected it? If a sample had a dense .text section, a
normal-looking import table and no known packer signature, what would you check
next? Why does a smaller window lower the maximum entropy you can observe?
Key takeaways
- Packed samples hide their real code; static triage of a packed file mostly describes the stub, so detect packing early.
- Entropy runs from 0 to 8 bits per byte; code sits around 5–6.5, compressed or encrypted data above about 7.2.
- Measure per section and with a sliding window, not just for the whole file.
- A verdict needs several indicators: dense section plus raw-size anomalies, writable-executable sections, entry point in an unusual section, a tiny import table or a packer signature.
- Compressed resources, installers and legitimate protectors also produce high entropy — judge by where the dense data is and whether it is executed.