Skip to content

Lesson 9.6 · Evasion & Unpacking· 50 min

Manual Unpacking to the OEP

Find the OEP with breakpoint tricks, dump the process, and rebuild a destroyed or hashed IAT into a binary that loads and runs like the original.

Objectives

  • Locate the OEP in a live debugger with the ESP-after-pushad trick, a memory breakpoint on newly-executable memory, and breaks on VirtualAlloc/VirtualProtect for multi-stage stubs
  • Dump a process at the OEP and explain why the raw dump will not run without further repair
  • Rebuild a destroyed or hashed Import Address Table with Scylla, and know what it can and cannot resolve automatically
  • Fix section headers and characteristics so a dumped image loads and executes like the original file
  • Recognise when manual unpacking will not work — virtualised stubs, stolen bytes, multi-layer protectors — and name the alternative

How Packers Work narrowed unpacking down to one question: where and when does the real code exist in a runnable form? It promised that a concrete debugger procedure would follow for the three tractable cases — a compressor, a crypter that runs in-process, and a loader that re-launches or injects its payload. This is that procedure. You will find the OEP, dump the process, and rebuild enough of the file to hand a working binary to the rest of your toolchain.

Nothing here defeats a real protector, and it is not meant to. The goal is the large middle of what you will actually meet: compressors, most crypters, and loaders that keep the payload in one piece somewhere you can reach.

Locating the OEP

The previous lesson found the tail jump by reading a disassembly listing end to end. In a live debugger you do not scroll — you let the stub's own behaviour trigger a stop exactly at the moment you want.

The ESP-after-pushad trick

Many simple stubs open by saving every general-purpose register with a single pushad-equivalent sequence (or, on x64, a run of individual push instructions) and close by restoring them symmetrically just before the tail jump. The address of the stack pointer immediately after that opening save is therefore the exact address it will return to right before the stub hands off control — the restores pop the same number of slots the pushes wrote.

  1. Step to the first instruction after the register-save sequence and read ESP/RSP.
  2. Set a hardware breakpoint on that stack address (memory, read or access, 4 or 8 bytes). x64dbg: right-click the address in the stack view → Set Hardware Access Breakpoint.
  3. Run. The next thing that reads that exact slot is the final pop that restores the last saved register — one or two instructions before the tail jump.

This works because a stub author writes the save/restore pair to leave the process looking untouched; that symmetry is what you are exploiting. It fails silently against a stub that saves registers a different way each time it runs, which is rare in compressors and common in protectors.

Memory breakpoints on write-then-execute

How Packers Work described the write-then-execute pattern: a region starts writable so the stub can fill it, then becomes executable so the CPU can run it. You can break on exactly that transition instead of guessing at stack symmetry:

  1. Let the stub run until it has allocated and started writing to the region that will hold the restored image (watch the memory map in x64dbg, or step past the first large VirtualAlloc/VirtualProtect call).
  2. Set a hardware execute breakpoint on the start of that region.
  3. Run. Execution stops on the very first instruction the restored code runs — which, if the stub jumps straight to the OEP, is the OEP.

This is more general than the ESP trick because it does not depend on how the stub manages the stack; it depends only on the write-then-execute shape, which almost every packer has.

Breaking on stage transitions

A multi-layer packer or a loader that decrypts a second executable moves through several VirtualAlloc / VirtualProtect / VirtualFree calls, one pair per stage. Set breakpoints on all three and log the region and size argument at each hit instead of the address alone:

CallWhat to recordWhy
VirtualAllocReturn address (base), dwSize, flProtectNew region appearing — likely the next stage's memory
VirtualProtectlpAddress, flNewProtectA region flipping to executable — the write-then-execute transition
VirtualFreelpAddressA stage being discarded — you may have already missed it

Each time VirtualProtect flips a region to PAGE_EXECUTE*, ask whether this is the final stage or an intermediate one — check whether the freshly executable bytes look like a real function prologue and a plausible import table, or like more stub code. If it is another stub, keep going; if it looks like a finished program, you have very likely found the OEP.

Tip: If the payload never runs in this process at all — because the stub calls CreateProcess, WriteProcessMemory and ResumeThread, as covered in How Packers Work — none of the three approaches above will find an OEP here. Follow the new process instead and repeat this section there.

Dumping the process

Once execution is sitting on the OEP, the process's memory holds a complete, correctly relocated copy of the original image — but a raw memory dump is not a runnable file. Two things differ from a normal PE on disk:

  • Section alignment. In memory, sections are aligned and sized to SectionAlignment (usually a 4 KB page); on disk they use the much smaller FileAlignment. A tool must convert from one to the other, matching raw offsets and raw sizes to what the section actually occupies.
  • The entry point field. The packed file's header still points at the stub's entry point, not the OEP you just found. A dumped image with an unchanged header would start executing the packer's cleanup code again, not the program.

Scylla (bundled with x64dbg, also standalone) and pe-sieve both perform this conversion. In x64dbg with Scylla attached: Attach to the process, confirm the entry point field shows the address you are stopped at, then Dump to write the fixed-alignment image to disk. pe-sieve does the equivalent from the command line and is useful when you want to script a batch of samples rather than drive a GUI for each one.

A dump taken at the OEP is not yet a working executable. It still has one broken piece: the import table.

Rebuilding the Import Address Table

How Packers Work explained why the stub destroys imports in the first place: the loader cannot read a packed import directory, so the stub resolves every function itself at runtime with LoadLibrary and GetProcAddress, then throws away the information that would let a normal loader do the same later. This is IAT destruction — sometimes deliberate obfuscation, sometimes just a side effect of how the stub works. Either way, the dump you just took has real, resolved function pointers sitting in memory with no names attached, and no import directory pointing at them.

Scylla's IAT rebuilding solves this by working backwards from the pointers themselves:

  1. Search the dumped process for the region that looks like an IAT: a run of pointer-sized values, each one landing inside a loaded module's code or export range. Scylla scans memory for this pattern and reports a candidate base address and size — confirm it against what you saw the stub build, or accept its guess and adjust the count if it is obviously wrong.
  2. Get imports. For every candidate slot, Scylla reads the pointer, finds which loaded module's address range contains it, and searches that module's export table for a matching address. A match becomes a named import (kernel32.dll!CreateFileW); no match is flagged invalid and left for you to resolve or discard.
  3. Fix dump. Scylla writes a new import directory into the dumped file, pointing at a fresh IAT built from the resolved names, and patches the file's OPTIONAL_HEADER.AddressOfEntryPoint to the OEP.

This works cleanly against a stub that resolves functions with ordinary GetProcAddress calls, because the resulting pointers really do point at each DLL's real exported code — the lookup in step 2 is exact.

It breaks down against API hashing. A hashing stub never calls GetProcAddress with a name; it computes a hash of each function name at build time, walks the target DLL's export directory at runtime, hashes each exported name the same way, and takes the pointer whose hash matches. The pointers Scylla finds are just as real and just as resolvable by address — step 1 and the "which module" half of step 2 still work — but there is no GetProcAddress call site to read a name from, so an automatic tool has nothing named to attach. You either recognise the hash algorithm (it is usually a small, distinctive routine visible right where the resolution happens) and build a lookup table offline, or you accept an IAT of correct-but-unnamed function pointers and identify calls by address and behaviour instead of by name as you continue analysis.

Warning: Always inspect Scylla's "invalid" list before trusting a fixed dump. A pointer that resolves to the middle of a function rather than its start, or into a security-product hook rather than the original code (see Hooking and User-Mode Rootkits), is a sign the search picked up the wrong region — rerun with a narrower base address and size before accepting the result.

Fixing headers and validating the result

Scylla's "Fix Dump" step, mentioned above, also repairs the section table entries the dump needed: raw offsets and sizes converted from the in-memory layout, and characteristics flags corrected if the packer had left them writable-and-executable. PE-bear is the fastest way to eyeball the result afterwards — open the fixed dump and confirm the entry point RVA now lands inside a real code section (not a packer-named one), that the import directory resolves to readable names, and that the section table's raw sizes are no longer zero anywhere they matter.

The only test that actually matters, though, is whether the file runs and behaves like the original. Run it in the isolated lab from Building a Safe Analysis Lab, compare its observable behaviour (network requests, files touched, output) against whatever you already know about the sample, and only then treat the fixed dump as the working artefact for the rest of your analysis — disassembly, strings, or handing it to a sandbox for a second opinion.

When this does not work

Three shapes from the previous lesson defeat the procedure above outright, and recognising them early saves you from chasing an OEP that is not there:

ObstacleWhat you will observeWhat to do instead
Virtualised or heavily mutated stub (e.g. Themida)No point where original x86 code appears in memory intact; the "write-then-execute" region holds a bytecode interpreter, not the programBehavioural analysis, or dedicated devirtualisation research — not a dump
Stolen bytesA dump at the "OEP" runs but crashes almost immediately, or produces obviously wrong output for its first few instructionsRecover the stolen instructions from inside the stub and patch them back before the tail-jump target
TLS callbacks doing the real workThe interesting behaviour already happened before your first breakpoint could fireEnumerate and break on TLS callbacks specifically, as covered in Anti-Debugging, before chasing an OEP at all

A multi-layer packer is not on this list because it is not a hard stop — it just means repeating this entire lesson once per layer, re-triaging after each dump the way How Packers Work described.

Lab: unpack your own UPX-packed program

This lab packs a harmless program you compile yourself, finds the OEP the way the previous lesson taught, and then goes one step further: rebuilding a runnable file from the dump and proving it behaves identically to the original. You need mingw-w64, upx, Wine (to run the Windows binaries), and a Python venv with pefile and capstone. Listings come from MinGW-w64 GCC 15.2.0, UPX 5.2.1, pefile 2024.8.26, Capstone 5.0.7 and Wine 11.0.

  1. Build and pack, and confirm the packed copy still runs correctly:

    bash
    python3 -m venv venv && ./venv/bin/pip install pefile capstone
    printf '#include <stdio.h>\nint main(void){ puts("hello, analyst"); return 0; }\n' > hello.c
    x86_64-w64-mingw32-gcc -O2 -s -o hello.exe hello.c
    upx -q -o hello_upx.exe hello.exe
    wine hello.exe
    wine hello_upx.exe

    Both print hello, analyst. From the outside, packed and unpacked are behaviourally identical — which is the whole point of a compressor, and exactly why static indicators (entropy, tiny import count) matter more than behaviour for detecting packing in the first place.

  2. Find the OEP, reusing the tail-jump search from the previous lesson. Save findoep.py:

    python
    # findoep.py — disassemble a packed file's stub from its entry point and
    # report the direct jmp whose target equals the original file's entry point.
    import sys, pefile
    from capstone import Cs, CS_ARCH_X86, CS_MODE_64
    
    orig_ep = pefile.PE(sys.argv[1]).OPTIONAL_HEADER.AddressOfEntryPoint
    pe = pefile.PE(sys.argv[2])
    ep = pe.OPTIONAL_HEADER.AddressOfEntryPoint
    sec = next(s for s in pe.sections
               if s.VirtualAddress <= ep < s.VirtualAddress + s.Misc_VirtualSize)
    code = sec.get_data()[ep - sec.VirtualAddress:]
    
    md = Cs(CS_ARCH_X86, CS_MODE_64)
    oep_rva = None
    for i in md.disasm(code, ep):
        if i.mnemonic == "jmp" and i.op_str.startswith("0x"):
            tgt = int(i.op_str, 16)
            if tgt == orig_ep:
                oep_rva = tgt
                print(f"tail jump at {i.address:#x}  ->  OEP at {tgt:#x}")
                break
    
    assert oep_rva == orig_ep, "did not find the tail jump"
    bash
    ./venv/bin/python findoep.py hello.exe hello_upx.exe
    text
    tail jump at 0xc9ab  ->  OEP at 0x1440

    In a live x64dbg session on the Windows VM, this is the address you would reach with the ESP-after-pushad trick or a hardware execute breakpoint on the newly-decompressed region, rather than by disassembling the file statically — confirm you understand why both routes land on the same instruction before moving on.

  3. Dump and rebuild, and check what changed. Rather than drive Scylla by hand for a UPX file specifically, use UPX's own -d switch — the same decompression algorithm a manual dump-and-fix would have to reimplement — and compare the section table before and after with inspect.py from the previous lesson:

    bash
    cp hello_upx.exe hello_dumped.exe
    upx -d -q hello_dumped.exe
    python
    # sections.py — compare entry point and import count across two files
    import sys, pefile
    for p in sys.argv[1:]:
        pe = pefile.PE(p)
        ep = pe.OPTIONAL_HEADER.AddressOfEntryPoint
        named = [i.name.decode() for d in getattr(pe, "DIRECTORY_ENTRY_IMPORT", [])
                 for i in d.imports if i.name]
        print(f"{p:16} entry_point={ep:#06x}  imports={len(named)}")
    bash
    ./venv/bin/python sections.py hello.exe hello_upx.exe hello_dumped.exe
    text
    hello.exe        entry_point=0x1440  imports=42
    hello_upx.exe    entry_point=0xc750  imports=12
    hello_dumped.exe entry_point=0x1440  imports=42

    The repaired file's entry point now matches the OEP you found in step 2 exactly, and the import count is back to the original 42 — the same outcome Scylla's "Get Imports" + "Fix Dump" sequence produces by walking memory pointers rather than replaying a known algorithm. This is the validation step: a correct manual unpack should always reproduce this entry-point-and-import-count match.

  4. Prove it behaves like the original, the final check from the lesson above:

    bash
    wine hello_dumped.exe
    text
    hello, analyst

    Identical output to both the original and the packed version. For this lab that is a formality; for a real sample it is the difference between a usable dump and a plausible-looking one that crashes the moment you hand it to a disassembler.

Questions to answer: If hello_upx.exe had used a custom XOR-based resolver instead of calling GetProcAddress for each import — the API-hashing case — which of the numbers in step 3's table would still match, and which would not? Why does the ESP-after-pushad trick still work even though this lab found the OEP by static disassembly instead? If step 3's import count had come back as 40 instead of 42, what would you check first in Scylla's invalid-import list? Sketch, in one sentence each, what changes about this whole procedure if the packer were a process-hollowing loader instead of a compressor.

Key takeaways

  • Find the OEP with a hardware breakpoint on the stack address exposed right after the stub's opening register save (ESP-after-pushad), or an execute breakpoint on the region that just became writable-then-executable; for multi-stage stubs, break on VirtualAlloc/VirtualProtect/VirtualFree and inspect each newly-executable region in turn.
  • A dump taken at the OEP is not runnable yet: in-memory section alignment does not match on-disk alignment, and the header's entry point field still points at the stub.
  • Scylla and pe-sieve convert alignment and rewrite the entry point; Scylla's IAT search then finds pointer-shaped memory, resolves each pointer to a named export by address, and writes a fresh import directory.
  • IAT rebuilding resolves pointers correctly even under API hashing, but cannot recover names the stub never looked up by name — expect an IAT of correct, unnamed functions in that case.
  • Validate a fixed dump by checking the entry point and import count against what you expect, then by running it and comparing behaviour to the original — not by assuming the dump worked because it opened in a disassembler.
  • Virtualised stubs, stolen bytes and TLS-callback-driven logic defeat this procedure outright; recognise them from the section table and imports before you invest time chasing an OEP that either does not exist or is not enough.