Skip to content

Leçon 3.4 · Triage statique· 35 min

Detecting Packers and Entropy

Recognise packed and encrypted executables with entropy, section anomalies, import tables and signatures — and avoid the classic false positives.

Cette leçon n’est disponible qu’en anglais pour le moment.

Objectifs

  • Explain what packers, crypters and protectors do and why they defeat static analysis
  • Compute Shannon entropy and interpret whole-file, per-section and sliding-window values
  • Combine entropy with structural indicators to reach a confident packing verdict
  • Recognise legitimate sources of high entropy that are not packing

Everything in this module so far — strings, imports, capabilities — assumes the interesting code is sitting in the file where you can read it. A packed sample breaks that assumption. Its real code is compressed or encrypted, and a small stub restores it in memory at runtime. Static triage of a packed file mostly describes the packer, not the malware.

So one of the first questions of any triage is: is this packed? This lesson teaches you to answer it with evidence, and to avoid the false alarms that trap beginners.

Packers, crypters and protectors

The terms overlap, but they point at different goals:

KindMain goalExamplesTypical stub behaviour
Packer (compressor)Smaller filesUPX, MPRESS, ASPackDecompress sections, rebuild imports, jump to the original entry point
CrypterEvade signature detectionCommodity crypters sold with malwareDecrypt an embedded payload, often into a new process or fresh memory
ProtectorResist reversing and tamperingThemida, VMProtect, EnigmaAnti-debugging, import obfuscation, code virtualization

All three share a runtime pattern:

text
  on disk                              in memory at runtime
  ┌───────────────┐                    ┌────────────────────┐
  │ headers       │                    │ headers            │
  │ stub code     │ ── stub runs ───►  │ stub code          │
  │ packed blob   │    decompresses /  │ original code  ◄───┼── jump to OEP
  │ (high entropy)│    decrypts,       │ original data      │
  └───────────────┘    fixes imports   │ rebuilt IAT        │
                                       └────────────────────┘

The stub ends by transferring control to the original entry point (OEP) of the unpacked code. Finding that moment and dumping memory is the job of the upcoming Unpacking module; this lesson is only about detection. The UPX packing and runtime crypter technique pages describe both families in more depth.

Entropy in one formula

Shannon entropy measures how unpredictable a block of bytes is. For a buffer where byte value i appears with probability pᵢ:

text
H = − Σ pᵢ · log₂(pᵢ)        summed over the 256 possible byte values

The result is in bits per byte, from 0.0 to 8.0:

  • 0.0 — every byte is identical (a block of zeros).
  • ≈ 3–5 — text, tables, structured data.
  • ≈ 5–6.5 — typical compiled x86/x64 code.
  • ≈ 7.2 and above — compressed or encrypted data; truly random bytes approach 8.0.

Compression and encryption both remove redundancy, which is why they push entropy toward the maximum. Plain code cannot get there: opcodes, registers and common instruction patterns repeat too much.

Tip: Entropy depends on how many bytes you measure. A 1 KiB window of perfectly random data only scores about 7.8, because 1024 samples cannot fill 256 buckets evenly. Compare values measured with the same window size.

Three ways to measure

  • Whole file — one number; quick but coarse. A large, low-entropy section can hide a small encrypted blob.
  • Per section — the most useful triage view: which part of the file is dense?
  • Sliding window — entropy for each consecutive block (say 1–4 KiB), plotted as a graph. Detect It Easy's entropy view and the lab script below both do this. It reveals blobs inside a section and data in the overlay.

Indicators beyond entropy

Entropy alone is never a verdict. Combine it with the structural indicators you learned in PE Sections and Reading Capabilities from Imports:

IndicatorWhy it suggests packing
A section with raw size 0 but a large virtual sizeSpace reserved for code that will be written at runtime (UPX's UPX0)
Writable + executable sectionThe stub must write the code it will later run
Entry point in the last section or outside .textExecution starts in the stub, not the program
Tiny import table with LoadLibrary, GetProcAddress, VirtualProtect / VirtualAllocImports are rebuilt at runtime (see dynamic import resolution)
Unusual section names (UPX0, .aspack, .themida, random strings)Packer-generated layout
Very few readable stringsStrings are inside the compressed blob
High-entropy overlay or resourcePayload stored outside the sections (see Resources and Overlays)

Three or more of these together, plus a dense section, is a confident verdict.

Signature-based identification

Tools such as Detect It Easy (DiE) match known byte patterns at the entry point and structural rules to name the packer, compiler and linker. That is fast and often exactly right for common packers — but signatures can be spoofed (a sample can include fake UPX section names) and custom crypters have no signature at all. Treat a DiE label as a lead to confirm, and treat "no packer detected" as "no known packer detected". capa likewise reports when a file looks packed and matches rules such as "packed with UPX".

The false positives

High entropy is common in perfectly legitimate files:

  • Compressed resources. PNG and JPEG icons, embedded ZIPs and fonts are already compressed. An .rsrc section full of images scores high.
  • Installers and self-extractors. NSIS, Inno Setup and 7-Zip SFX files carry a compressed archive in the overlay — by design.
  • Legitimate protectors. Commercial software, games and licensed tools often use protectors; .NET applications may be obfuscated.
  • Embedded certificates and crypto tables — small but dense.

The discriminator is where the entropy is. Dense data in a resource that the program loads as an image is normal. Dense data in a section the entry point lives in, or a blob that is decrypted into executable memory, is not.

Lab: measure, then judge

You need mingw-w64, UPX and Python 3 with pefile. You will compare three harmless binaries: a plain program, the same program packed with UPX, and a program that embeds 32 KiB of random bytes as a resource.

  1. Build the three variants:

    bash
    printf '#include <stdio.h>\nint main(void){puts("hello, analyst");return 0;}\n' > hello.c
    x86_64-w64-mingw32-gcc -s -o hello.exe hello.c
    upx -o hello_upx.exe hello.exe
    
    head -c 32768 /dev/urandom > blob.bin
    echo '101 RCDATA "blob.bin"' > blob.rc
    x86_64-w64-mingw32-windres blob.rc -O coff -o blob.res.o
    x86_64-w64-mingw32-gcc -s -o hello_blob.exe hello.c blob.res.o
  2. Save this script as entropy.py:

    python
    # entropy.py — per-section and sliding-window entropy for a PE file
    import math, sys
    from collections import Counter
    import pefile
    
    def entropy(buf: bytes) -> float:
        """Shannon entropy in bits per byte, 0.0 to 8.0."""
        n = len(buf)
        if n == 0:
            return 0.0
        return max(0.0, -sum(c / n * math.log2(c / n) for c in Counter(buf).values()))
    
    def region(pe, off: int) -> str:
        """Name the part of the file that contains file offset `off`."""
        for s in pe.sections:
            if s.PointerToRawData <= off < s.PointerToRawData + s.SizeOfRawData:
                return s.Name.rstrip(b"\x00").decode(errors="replace")
        end = max(s.PointerToRawData + s.SizeOfRawData for s in pe.sections)
        return "overlay" if off >= end else "headers"
    
    path = sys.argv[1]
    window = int(sys.argv[2]) if len(sys.argv) > 2 else 1024
    data = open(path, "rb").read()
    pe = pefile.PE(data=data)
    
    print(f"{path}: {len(data)} bytes, whole-file entropy {entropy(data):.2f}\n")
    print(f"{'section':8} {'raw size':>9} {'entropy':>8}")
    for s in pe.sections:
        name = s.Name.rstrip(b"\x00").decode(errors="replace")
        print(f"{name:8} {s.SizeOfRawData:>9} {entropy(s.get_data()):>8.2f}")
    
    print(f"\nwindow = {window} bytes; one '#' = 0.25 bits/byte; '|' marks 7.2")
    for off in range(0, len(data), window):
        h = entropy(data[off:off + window])
        bar = list(("#" * round(h * 4)).ljust(32))
        bar[29] = "|"
        print(f"{off:#08x} {region(pe, off):8} {h:4.2f} {''.join(bar)}")
  3. Run it on the plain program:

    bash
    python3 entropy.py hello.exe 4096
    text
    hello.exe: 16384 bytes, whole-file entropy 4.70
    
    section   raw size  entropy
    .text         7168     5.69
    .data          512     0.57
    .rdata        2560     4.02
    .pdata        1024     2.30
    .xdata         512     3.31
    .bss             0     0.00
    .idata        2560     3.52
    .tls           512     0.00
    .reloc         512     1.24

    Code sits in the expected 5–6.5 band; nothing crosses 7.2.

  4. Run it on the UPX build:

    text
    hello_upx.exe: 8704 bytes, whole-file entropy 7.00
    
    section   raw size  entropy
    UPX0             0     0.00
    UPX1          7168     7.48
    UPX2          1024     3.31

    Note every structural indicator at once: UPX0 has no raw data, UPX1 is dense, the section names are a signature, and the import table (run capimports.py from the previous lesson) is reduced to LoadLibraryA, GetProcAddress, VirtualProtect and ExitProcess plus one import per C runtime DLL. In the sliding-window view, the first window is labelled headers but scores high: with a 4096-byte window it already spans most of UPX1. Re-run with 1024 to see the boundary sharply.

  5. Run it on the program with the random resource:

    text
    hello_blob.exe: 49664 bytes, whole-file entropy 7.27
    ...
    .text         7168     5.69
    .rsrc        33280     7.97

    The whole-file entropy is higher than the UPX build, yet this program is not packed at all: the code is untouched, the imports are normal, and the dense bytes sit in a resource the program never executes.

  6. Optional: open all three in Detect It Easy (entropy graph and signature scan) and in PE-bear, and compare with your script.

Questions to answer: Which single number would have misled you in step 5, and which indicators corrected it? If a sample had a dense .text section, a normal-looking import table and no known packer signature, what would you check next? Why does a smaller window lower the maximum entropy you can observe?

Key takeaways

  • Packed samples hide their real code; static triage of a packed file mostly describes the stub, so detect packing early.
  • Entropy runs from 0 to 8 bits per byte; code sits around 5–6.5, compressed or encrypted data above about 7.2.
  • Measure per section and with a sliding window, not just for the whole file.
  • A verdict needs several indicators: dense section plus raw-size anomalies, writable-executable sections, entry point in an unusual section, a tiny import table or a packer signature.
  • Compressed resources, installers and legitimate protectors also produce high entropy — judge by where the dense data is and whether it is executed.