Skip to content

Lesson 9.5 · Evasion & Unpacking· 45 min

How Packers Work

What a packer's unpacking stub does step by step — decompress, rebuild imports, relocate, jump to the OEP — and what it means for unpacking.

Objectives

  • Describe what an unpacking stub does at runtime, from decompressing sections to transferring control to the original entry point
  • Read the section table and memory map of a packed file and predict how they change once the stub has run
  • Recognise the tail jump and the other shapes a stub uses to reach the OEP
  • Distinguish compressors, crypters and protectors by what their stubs do to code and imports
  • Explain how multi-stage and process-injecting crypters change where the payload ends up, and what each family implies for unpacking

Detecting Packers and Entropy taught you to recognise a packed file, and Dumping Memory and Extracting Payloads taught you to copy the unpacked result out of a live process and repair its layout. This lesson fills the gap between them: what actually happens inside a packed executable between the loader handing it control and the real code running. Once you can narrate that sequence, the choices you make when unpacking by hand — where to break, when to dump, what to rebuild — stop being a recipe and become obvious.

We are still doing detection and comprehension here, not unpacking. Nothing in this lesson runs a real sample; the lab packs a harmless program you built yourself.

The stub is a tiny loader

A packer takes a finished executable and produces a new one that carries the original as inert, transformed data plus a small program called the unpacking stub. The operating system loader knows nothing about any of this. It maps the packed file, reads the entry point from the PE header, and jumps there — straight into the stub. The stub then does, in user code, the parts of loading the OS already did for it and the parts the OS cannot do because the real image is compressed or encrypted.

Think of the stub as a second loader that runs inside the process. A thorough one performs these steps, roughly in order:

  1. Decompress or decrypt the packed blob into the memory the header reserved for it. Compressors run an algorithm like LZMA; crypters run a cipher, sometimes just a rolling XOR (see self-modifying code, which is what in-place decryption is).
  2. Place each original section at its RVA. The packed file's headers were crafted so the loader reserved a region the right size; the stub writes the restored bytes into it, section by section, the way Loading and Execution described the loader doing.
  3. Rebuild the import table. The original IAT is not filled in — the loader could not read a packed import directory. The stub walks the original import information and resolves every function itself, almost always with LoadLibrary and GetProcAddress (covered in Imports, Exports and the IAT).
  4. Apply base relocations. If the image did not load at its preferred base — which, with ASLR, it usually does not — every absolute address in the restored code is wrong. The stub reads the original relocation table and adds the load delta to each listed pointer.
  5. Handle TLS. If the original program used thread-local storage or TLS callbacks, the stub must restore the TLS directory so those callbacks fire. Malware likes TLS callbacks because they run before the entry point; a stub that forgets them breaks such samples.
  6. Restore or fix headers. Some packers repair the in-memory PE header so the unpacked image looks normal — section names, characteristics, the entry point field. Others leave the header packer-shaped, which is itself a tell.
  7. Transfer control to the original entry point (OEP). The last thing the stub does is jump to the address where the original program's execution began.

Not every packer does all seven. A compressor may skip anti-analysis work entirely; a protector may add integrity checks, anti-debugging and virtualisation around every step. But the skeleton is always "restore the image, wire up its imports, jump into it."

Tip: The imports a stub needs for its own work leak the whole scheme. A tiny import table containing LoadLibrary, GetProcAddress and VirtualProtect/VirtualAlloc says "this program resolves its real imports at runtime" — the exact indicator from the detection lesson, now explained.

Before and after, in the section table

The packed file on disk and the process after the stub has run are two different programs sharing a PE header. Reading both views side by side is the core skill of this lesson.

AspectPacked file, on diskAfter the stub runs, in memory
Entry pointInside the stub, typically in the last sectionExecution has moved to the OEP in the restored code
SectionsPacker-named (UPX0, UPX1, .aspack), one or more with raw size 0Original sections present at their RVAs, now populated
A large empty sectionSizeOfRawData = 0, big VirtualSize — reserved spaceFilled with the decompressed original code
ImportsTiny — just what the stub needsFull IAT, resolved to real API addresses
StringsFew; the rest are inside the blobThe program's real strings are readable
EntropyA dense, high-entropy sectionNormal code entropy where the blob was

The empty-section-becomes-code move is worth picturing concretely. UPX reserves UPX0 with no raw data at all; the compressed original lives in UPX1. At runtime the stub decompresses UPX1 into UPX0, so a section that was zero bytes on disk becomes the program's .text in memory. Every tool that trusts the on-disk section table is therefore looking at the wrong bytes until you dump and repair the image — which is precisely the layout problem from Dumping Memory.

The memory map tells the same story from the other side. In x64dbg's Memory Map or System Informer's Memory tab, a packed process shows a region that starts writable (so the stub can fill it) and becomes executable (so the CPU can run it) — the write-then-execute pattern of every unpacker.

The tail jump and its disguises

The instruction that transfers control from the stub to the OEP is called the tail jump, and finding it is how you locate the OEP for a dump. In the simplest packers it is exactly what the name says: a single jmp into the region the stub just filled, sitting near the end of the stub after a run of register restores.

Because it is so recognisable, packers disguise it. Watch for these shapes:

ShapeWhat you seeWhy it is used
Direct jumpjmp <oep> with a target in a different sectionSimplest; UPX and most compressors
ReturnPush the OEP, then retHides the target from a casual reader; see call/ret
Indirectjmp eax / call eax after the OEP is computed into a registerTarget is not a literal in the code
Via an OS callNtContinue/ZwContinue with a CONTEXT record whose Rip is the OEPControl transfer laundered through the kernel
Exception-drivenTrigger an exception; the handler resumes at the OEPAlso doubles as an anti-debug check

The reliable signal is not the opcode but the destination: control leaves the stub's section and lands in a region that was empty on disk and is now full of code, after which the register state looks like a fresh program (a normal prologue, not more stub bookkeeping). In the lab you will find a plain jmp whose target equals the original file's entry point exactly — the clean case that trains your eye for the disguised ones.

Compressors, crypters and protectors

The detection lesson introduced these three families by goal. Now compare them by what their stubs actually do, because that is what decides your unpacking effort.

Compressors exist to make files smaller. UPX is the archetype: a documented algorithm, a small predictable stub, a real tail jump, and — usefully — a built-in -d switch that reverses the transform. The original code is present in one piece after decompression, so a single well-timed dump gets you everything. See UPX packing.

Crypters exist to defeat signatures. The stub decrypts an embedded payload, often with a key that changes per build so no two samples share bytes. A crypter may still hand control to the payload with a tail jump in the same process — or it may not run the payload in this process at all (see the next section). The code is recoverable, but you must find the moment after decryption and before any wipe.

Protectors exist to resist reversing itself. They layer everything the other two families skip: anti-debugging (Anti-Debugging), anti-VM and sandbox checks, anti-disassembly (Anti-Disassembly), and code transformations that mean there may be no clean OEP to jump to at all. Three protector techniques change the game:

  • Virtualised or mutated stubs. VMProtect and Themida can compile the original code into bytecode for a custom virtual machine bundled in the stub, or mutate it into functionally equivalent but unreadable instructions. There is no moment when the original x86 appears in memory, so "dump at the OEP" does not apply. This is virtual-machine obfuscation, and defeating it is a research effort, not a dump.
  • Stolen bytes. The protector copies the first handful of instructions from the OEP into its own stub and executes them there, then jumps into the middle of the original function. A dump taken at the "OEP" is missing its first instructions; you must recover the stolen bytes from the stub and graft them back.
  • Import redirection. Instead of a normal IAT, each call goes through a stub that resolves the API at call time, often by API hashing. The IAT you dump points at packer thunks, not real functions, so import rebuilding — the separate repair from the memory-dumping lesson — becomes the hard part.
FamilyStub complexityImport handlingOriginal code in memory?Typical approach
CompressorSmallRebuilt in one passYes, intactOfficial unpacker or one dump
CrypterSmall–mediumRebuilt, or payload re-launchedYes, after decryptionBreak after decrypt, dump
ProtectorLargeRedirected/virtualisedSometimes neverCase by case; may need instrumentation

When the payload leaves the process

Not every stub jumps to an OEP in its own address space. Many commodity crypters are really loaders: they decrypt a second executable and run that, which changes where you go looking.

  • Multi-stage. Stage one decrypts and executes stage two, which may decrypt stage three. Each stage can use a different technique. You unpack one layer at a time, re-triaging after each, because the entropy, imports and even the file type can change from stage to stage.
  • Run from memory. The stub allocates memory, writes the payload PE, applies relocations and imports itself (a reflective loader), and calls its entry point. There is no tail jump into a packer section — the OEP is in freshly allocated private memory, which is exactly what the memory-dumping lesson taught you to hunt for.
  • Inject into another process. The stub spawns or opens a second process and writes the payload there — process hollowing replaces a suspended process's image; other techniques map code into a running one. The unpacked payload never executes in the file you launched, so dumping this process gives you the loader, not the malware. You dump the target process instead, at the moment control is handed to it.

This is why the memory-dumping lesson insisted on watching CreateProcess, WriteProcessMemory and ResumeThread: a stub that calls them is telling you the OEP is somewhere else.

What this means for unpacking

Everything above narrows to a single triage question you answer before you start unpacking: where and when does the real code exist in a runnable form?

  • A compressor with a known name: try its official unpacker first; if not, one dump at the tail jump.
  • A crypter that runs the payload in-process: break after decryption finishes, before it is used and possibly wiped, then dump and rebuild imports.
  • A loader that re-launches or injects: follow the new region or the new process, and dump there.
  • A protector: confirm whether the original code ever appears intact. If it does, the crypter workflow applies. If it is virtualised, dumping will not help and you move to behavioural analysis or dedicated devirtualisation tools.

An upcoming lesson on manual unpacking turns the first three of these into a concrete debugger procedure. The point of this one is that you can already predict which of them you are facing from the section table, the imports and the stub's API calls.

Lab: pack a program and find the tail jump

You will pack a harmless program with UPX, compare the two files' structure, and find the stub's jump to the OEP with capstone — confirming its target equals the original entry point. You need mingw-w64, upx (check with which upx), and a virtual environment with pefile and capstone. Listings below come from MinGW-w64 GCC, UPX 5.2.1, pefile 2024.8.26 and Capstone 5.0.7; your addresses will differ.

  1. Set up and build a plain program, then pack a copy:

    bash
    python3 -m venv venv && ./venv/bin/pip install pefile capstone
    printf '#include <stdio.h>\nint main(void){ puts("hello, analyst"); return 0; }\n' > hello.c
    x86_64-w64-mingw32-gcc -O2 -s -o hello.exe hello.c
    which upx && upx -q -o hello_upx.exe hello.exe
  2. Save this as inspect.py — it prints the entry point, the section table with permission flags, and the imports for any PE:

    python
    # inspect.py — section table, entry point and imports for a PE file
    import sys, pefile
    
    for p in sys.argv[1:]:
        pe = pefile.PE(p)
        oh = pe.OPTIONAL_HEADER
        print(f"== {p}  ep_rva={oh.AddressOfEntryPoint:#x}")
        print(f"   {'name':8} {'VSize':>8} {'RawSize':>8} flags")
        for s in pe.sections:
            n = s.Name.rstrip(b"\x00").decode(errors="replace")
            c = s.Characteristics
            f = ("X" if c & 0x20000000 else "") + ("R" if c & 0x40000000 else "") \
                + ("W" if c & 0x80000000 else "")
            print(f"   {n:8} {s.Misc_VirtualSize:>#8x} {s.SizeOfRawData:>#8x} {f}")
        named = [i.name.decode() for d in getattr(pe, "DIRECTORY_ENTRY_IMPORT", [])
                 for i in d.imports if i.name]
        print(f"   imports: {len(named)} named function(s)")
        ep = oh.AddressOfEntryPoint
        sec = next(s.Name.rstrip(b"\x00").decode() for s in pe.sections
                   if s.VirtualAddress <= ep < s.VirtualAddress + max(s.Misc_VirtualSize, s.SizeOfRawData))
        print(f"   entry point is in section {sec}\n")
    bash
    ./venv/bin/python inspect.py hello.exe hello_upx.exe
    text
    == hello.exe  ep_rva=0x1440
       name        VSize  RawSize flags
       .text      0x1a20   0x1c00 XR
       .data        0xa0    0x200 RW
       .rdata      0x9c8    0xa00 R
       ...
       .reloc       0x60    0x200 R
       imports: 42 named function(s)
       entry point is in section .text
    
    == hello_upx.exe  ep_rva=0xc750
       name        VSize  RawSize flags
       UPX0       0xa000    0x000 XRW
       UPX1       0x2000   0x1c00 XRW
       UPX2       0x1000    0x400 RW
       imports: 12 named function(s)
       entry point is in section UPX1

    Read every indicator from the detection lesson at once: UPX0 has a large virtual size but zero raw data — the space the stub will fill; the sections are writable and executable; the entry point sits in UPX1, the stub, not in the original code; and the import count collapsed from 42 to 12. The four KERNEL32.DLL imports UPX keeps are LoadLibraryA, GetProcAddress, VirtualProtect and ExitProcess — the toolkit for rebuilding the rest at runtime.

  3. Save this as tailjmp.py. It disassembles the stub from the packed entry point and reports every direct jmp, flagging any whose target equals the original file's entry point:

    python
    # tailjmp.py — find the UPX stub's jump to the original entry point
    import pefile
    from capstone import Cs, CS_ARCH_X86, CS_MODE_64
    
    orig_ep = pefile.PE("hello.exe").OPTIONAL_HEADER.AddressOfEntryPoint
    pe = pefile.PE("hello_upx.exe")
    ep = pe.OPTIONAL_HEADER.AddressOfEntryPoint
    sec = next(s for s in pe.sections
               if s.VirtualAddress <= ep < s.VirtualAddress + s.Misc_VirtualSize)
    code = sec.get_data()[ep - sec.VirtualAddress:]
    
    md = Cs(CS_ARCH_X86, CS_MODE_64)
    insns = list(md.disasm(code, ep))
    print(f"original OEP rva = {orig_ep:#x}; packed entry rva = {ep:#x}\n")
    print("first stub instructions:")
    for i in insns[:4]:
        print(f"  {i.address:#08x}  {i.mnemonic:5} {i.op_str}")
    print("\ndirect jumps in the stub:")
    for i in insns:
        if i.mnemonic == "jmp" and i.op_str.startswith("0x"):
            tgt = int(i.op_str, 16)
            mark = "  <== target == original OEP" if tgt == orig_ep else ""
            print(f"  {i.address:#08x}  jmp {tgt:#x}{mark}")
    bash
    ./venv/bin/python tailjmp.py
    text
    original OEP rva = 0x1440; packed entry rva = 0xc750
    
    first stub instructions:
      0x00c750  push  rbx
      0x00c751  push  rsi
      0x00c752  push  rdi
      0x00c753  push  rbp
    
    direct jumps in the stub:
      0x00c7c3  jmp 0xc7cd
      0x00c844  jmp 0xc7cd
      ...
      0x00c9ab  jmp 0x1440  <== target == original OEP
      0x00c9d2  jmp 0xc9b9

    The stub opens by saving registers — it intends to restore them and leave the process looking untouched. Most of its jumps stay inside itself (the decompression loops). Exactly one leaves for 0x1440, the address inspect.py reported as hello.exe's entry point. That is the tail jump, and its target is the OEP.

  4. Look at the instructions right before it. Re-run with a slice around the tail jump (edit tailjmp.py to print insns near that address), and you will see the register restores and stack fixup that mark the end of a stub:

    text
    0x00c999  pop   rsi
    0x00c99a  pop   rbx
    0x00c9a7  sub   rsp, -0x80
    0x00c9ab  jmp   0x1440      ; e9 90 4a ff ff
    0x00c9b0  ret

    The pops undo the pushes from step 3, the stack is squared away, and only then does control leave for the OEP. The encoding e9 90 4a ff ff is a near jmp with a 32-bit relative displacement — worth recognising by sight.

Questions to answer: Why does UPX0 have a raw size of zero on disk, and what will occupy that region after the stub runs? If you dumped this process the instant the tail jump executed and trusted the on-disk section table, which bytes would your disassembler show at the "entry point", and why? The stub keeps only four KERNEL32 imports — which step of the stub's job needs each of them? If a packer replaced the jmp 0x1440 with push 0x1440; ret, how would you still identify the OEP?

Key takeaways

  • The unpacking stub is a second loader inside the process: it decompresses or decrypts the original image, places sections at their RVAs, rebuilds the IAT, applies relocations, handles TLS, may fix headers, and finally jumps to the OEP.
  • The packed file on disk and the unpacked process share one header but are different programs; an empty section on disk (UPX0) becomes the real code in memory, which is why dumps need layout repair.
  • The tail jump transfers control from the stub to the OEP. Its opcode varies — jmp, ret, indirect, NtContinue, exception — but its destination is always the freshly restored code.
  • Compressors keep the code intact and are easiest; crypters decrypt a payload and may re-launch it; protectors add anti-analysis and can virtualise or steal code so no clean OEP exists.
  • Multi-stage, reflective and injecting stubs move the payload into new memory or another process, so the OEP — and the dump you want — may not be in the file you ran.
  • Decide where and when the real code exists in runnable form before unpacking; the section table, imports and the stub's API calls tell you which family you face.