Lesson 9.6 · Evasion & Unpacking· 50 min
Manual Unpacking to the OEP
Find the OEP with breakpoint tricks, dump the process, and rebuild a destroyed or hashed IAT into a binary that loads and runs like the original.
Objectives
- Locate the OEP in a live debugger with the ESP-after-pushad trick, a memory breakpoint on newly-executable memory, and breaks on VirtualAlloc/VirtualProtect for multi-stage stubs
- Dump a process at the OEP and explain why the raw dump will not run without further repair
- Rebuild a destroyed or hashed Import Address Table with Scylla, and know what it can and cannot resolve automatically
- Fix section headers and characteristics so a dumped image loads and executes like the original file
- Recognise when manual unpacking will not work — virtualised stubs, stolen bytes, multi-layer protectors — and name the alternative
How Packers Work narrowed unpacking down to one question: where and when does the real code exist in a runnable form? It promised that a concrete debugger procedure would follow for the three tractable cases — a compressor, a crypter that runs in-process, and a loader that re-launches or injects its payload. This is that procedure. You will find the OEP, dump the process, and rebuild enough of the file to hand a working binary to the rest of your toolchain.
Nothing here defeats a real protector, and it is not meant to. The goal is the large middle of what you will actually meet: compressors, most crypters, and loaders that keep the payload in one piece somewhere you can reach.
Locating the OEP
The previous lesson found the tail jump by reading a disassembly listing end to end. In a live debugger you do not scroll — you let the stub's own behaviour trigger a stop exactly at the moment you want.
The ESP-after-pushad trick
Many simple stubs open by saving every general-purpose register with a single
pushad-equivalent sequence (or, on x64, a run of individual
push instructions) and close by restoring them
symmetrically just before the tail jump. The address of the stack pointer
immediately after that opening save is therefore the exact address it will
return to right before the stub hands off control — the restores pop the same
number of slots the pushes wrote.
- Step to the first instruction after the register-save sequence and read
ESP/RSP. - Set a hardware breakpoint on that stack address (memory, read or access, 4 or 8 bytes). x64dbg: right-click the address in the stack view → Set Hardware Access Breakpoint.
- Run. The next thing that reads that exact slot is the final
popthat restores the last saved register — one or two instructions before the tail jump.
This works because a stub author writes the save/restore pair to leave the process looking untouched; that symmetry is what you are exploiting. It fails silently against a stub that saves registers a different way each time it runs, which is rare in compressors and common in protectors.
Memory breakpoints on write-then-execute
How Packers Work described the write-then-execute pattern: a region starts writable so the stub can fill it, then becomes executable so the CPU can run it. You can break on exactly that transition instead of guessing at stack symmetry:
- Let the stub run until it has allocated and started writing to the region
that will hold the restored image (watch the memory map in x64dbg, or step
past the first large
VirtualAlloc/VirtualProtectcall). - Set a hardware execute breakpoint on the start of that region.
- Run. Execution stops on the very first instruction the restored code runs — which, if the stub jumps straight to the OEP, is the OEP.
This is more general than the ESP trick because it does not depend on how the stub manages the stack; it depends only on the write-then-execute shape, which almost every packer has.
Breaking on stage transitions
A multi-layer packer or a loader that
decrypts a second executable moves through several VirtualAlloc /
VirtualProtect / VirtualFree calls, one pair per stage. Set breakpoints on
all three and log the region and size argument at each hit instead of the
address alone:
| Call | What to record | Why |
|---|---|---|
VirtualAlloc | Return address (base), dwSize, flProtect | New region appearing — likely the next stage's memory |
VirtualProtect | lpAddress, flNewProtect | A region flipping to executable — the write-then-execute transition |
VirtualFree | lpAddress | A stage being discarded — you may have already missed it |
Each time VirtualProtect flips a region to PAGE_EXECUTE*, ask whether this
is the final stage or an intermediate one — check whether the freshly
executable bytes look like a real function prologue and a plausible import
table, or like more stub code. If it is another stub, keep going; if it looks
like a finished program, you have very likely found the OEP.
Tip: If the payload never runs in this process at all — because the stub calls
CreateProcess,WriteProcessMemoryandResumeThread, as covered in How Packers Work — none of the three approaches above will find an OEP here. Follow the new process instead and repeat this section there.
Dumping the process
Once execution is sitting on the OEP, the process's memory holds a complete, correctly relocated copy of the original image — but a raw memory dump is not a runnable file. Two things differ from a normal PE on disk:
- Section alignment. In memory, sections are aligned and sized to
SectionAlignment(usually a 4 KB page); on disk they use the much smallerFileAlignment. A tool must convert from one to the other, matching raw offsets and raw sizes to what the section actually occupies. - The entry point field. The packed file's header still points at the stub's entry point, not the OEP you just found. A dumped image with an unchanged header would start executing the packer's cleanup code again, not the program.
Scylla (bundled with x64dbg, also standalone) and pe-sieve both perform
this conversion. In x64dbg with Scylla attached: Attach to the process,
confirm the entry point field shows the address you are stopped at, then
Dump to write the fixed-alignment image to disk. pe-sieve does the
equivalent from the command line and is useful when you want to script a batch
of samples rather than drive a GUI for each one.
A dump taken at the OEP is not yet a working executable. It still has one broken piece: the import table.
Rebuilding the Import Address Table
How Packers Work explained why the stub destroys
imports in the first place: the loader cannot read a packed import directory,
so the stub resolves every function itself at runtime with LoadLibrary and
GetProcAddress, then throws away the information that would let a normal
loader do the same later. This is IAT destruction —
sometimes deliberate obfuscation, sometimes just a side effect of how the stub
works. Either way, the dump you just took has real, resolved function pointers
sitting in memory with no names attached, and no import directory pointing at
them.
Scylla's IAT rebuilding solves this by working backwards from the pointers themselves:
- Search the dumped process for the region that looks like an IAT: a run of pointer-sized values, each one landing inside a loaded module's code or export range. Scylla scans memory for this pattern and reports a candidate base address and size — confirm it against what you saw the stub build, or accept its guess and adjust the count if it is obviously wrong.
- Get imports. For every candidate slot, Scylla reads the pointer, finds
which loaded module's address range contains it, and searches that
module's export table for a matching address. A match becomes a named
import (
kernel32.dll!CreateFileW); no match is flagged invalid and left for you to resolve or discard. - Fix dump. Scylla writes a new import directory into the dumped file,
pointing at a fresh IAT built from the resolved names, and patches the
file's
OPTIONAL_HEADER.AddressOfEntryPointto the OEP.
This works cleanly against a stub that resolves functions with ordinary
GetProcAddress calls, because the resulting pointers really do point at each
DLL's real exported code — the lookup in step 2 is exact.
It breaks down against API hashing. A hashing
stub never calls GetProcAddress with a name; it computes a hash of each
function name at build time, walks the target DLL's export directory at
runtime, hashes each exported name the same way, and takes the pointer whose
hash matches. The pointers Scylla finds are just as real and just as
resolvable by address — step 1 and the "which module" half of step 2 still
work — but there is no GetProcAddress call site to read a name from, so an
automatic tool has nothing named to attach. You either recognise the hash
algorithm (it is usually a small, distinctive routine visible right where the
resolution happens) and build a lookup table offline, or you accept an IAT of
correct-but-unnamed function pointers and identify calls by address and
behaviour instead of by name as you continue analysis.
Warning: Always inspect Scylla's "invalid" list before trusting a fixed dump. A pointer that resolves to the middle of a function rather than its start, or into a security-product hook rather than the original code (see Hooking and User-Mode Rootkits), is a sign the search picked up the wrong region — rerun with a narrower base address and size before accepting the result.
Fixing headers and validating the result
Scylla's "Fix Dump" step, mentioned above, also repairs the section table entries the dump needed: raw offsets and sizes converted from the in-memory layout, and characteristics flags corrected if the packer had left them writable-and-executable. PE-bear is the fastest way to eyeball the result afterwards — open the fixed dump and confirm the entry point RVA now lands inside a real code section (not a packer-named one), that the import directory resolves to readable names, and that the section table's raw sizes are no longer zero anywhere they matter.
The only test that actually matters, though, is whether the file runs and behaves like the original. Run it in the isolated lab from Building a Safe Analysis Lab, compare its observable behaviour (network requests, files touched, output) against whatever you already know about the sample, and only then treat the fixed dump as the working artefact for the rest of your analysis — disassembly, strings, or handing it to a sandbox for a second opinion.
When this does not work
Three shapes from the previous lesson defeat the procedure above outright, and recognising them early saves you from chasing an OEP that is not there:
| Obstacle | What you will observe | What to do instead |
|---|---|---|
| Virtualised or heavily mutated stub (e.g. Themida) | No point where original x86 code appears in memory intact; the "write-then-execute" region holds a bytecode interpreter, not the program | Behavioural analysis, or dedicated devirtualisation research — not a dump |
| Stolen bytes | A dump at the "OEP" runs but crashes almost immediately, or produces obviously wrong output for its first few instructions | Recover the stolen instructions from inside the stub and patch them back before the tail-jump target |
| TLS callbacks doing the real work | The interesting behaviour already happened before your first breakpoint could fire | Enumerate and break on TLS callbacks specifically, as covered in Anti-Debugging, before chasing an OEP at all |
A multi-layer packer is not on this list because it is not a hard stop — it just means repeating this entire lesson once per layer, re-triaging after each dump the way How Packers Work described.
Lab: unpack your own UPX-packed program
This lab packs a harmless program you compile yourself, finds the OEP the way
the previous lesson taught, and then goes one step further: rebuilding a
runnable file from the dump and proving it behaves identically to the
original. You need mingw-w64, upx, Wine (to run the Windows binaries), and
a Python venv with pefile and capstone. Listings come from MinGW-w64 GCC
15.2.0, UPX 5.2.1, pefile 2024.8.26, Capstone 5.0.7 and Wine 11.0.
-
Build and pack, and confirm the packed copy still runs correctly:
bash python3 -m venv venv && ./venv/bin/pip install pefile capstone printf '#include <stdio.h>\nint main(void){ puts("hello, analyst"); return 0; }\n' > hello.c x86_64-w64-mingw32-gcc -O2 -s -o hello.exe hello.c upx -q -o hello_upx.exe hello.exe wine hello.exe wine hello_upx.exeBoth print
hello, analyst. From the outside, packed and unpacked are behaviourally identical — which is the whole point of a compressor, and exactly why static indicators (entropy, tiny import count) matter more than behaviour for detecting packing in the first place. -
Find the OEP, reusing the tail-jump search from the previous lesson. Save
findoep.py:python # findoep.py — disassemble a packed file's stub from its entry point and # report the direct jmp whose target equals the original file's entry point. import sys, pefile from capstone import Cs, CS_ARCH_X86, CS_MODE_64 orig_ep = pefile.PE(sys.argv[1]).OPTIONAL_HEADER.AddressOfEntryPoint pe = pefile.PE(sys.argv[2]) ep = pe.OPTIONAL_HEADER.AddressOfEntryPoint sec = next(s for s in pe.sections if s.VirtualAddress <= ep < s.VirtualAddress + s.Misc_VirtualSize) code = sec.get_data()[ep - sec.VirtualAddress:] md = Cs(CS_ARCH_X86, CS_MODE_64) oep_rva = None for i in md.disasm(code, ep): if i.mnemonic == "jmp" and i.op_str.startswith("0x"): tgt = int(i.op_str, 16) if tgt == orig_ep: oep_rva = tgt print(f"tail jump at {i.address:#x} -> OEP at {tgt:#x}") break assert oep_rva == orig_ep, "did not find the tail jump"bash ./venv/bin/python findoep.py hello.exe hello_upx.exetext tail jump at 0xc9ab -> OEP at 0x1440In a live x64dbg session on the Windows VM, this is the address you would reach with the ESP-after-pushad trick or a hardware execute breakpoint on the newly-decompressed region, rather than by disassembling the file statically — confirm you understand why both routes land on the same instruction before moving on.
-
Dump and rebuild, and check what changed. Rather than drive Scylla by hand for a UPX file specifically, use UPX's own
-dswitch — the same decompression algorithm a manual dump-and-fix would have to reimplement — and compare the section table before and after withinspect.pyfrom the previous lesson:bash cp hello_upx.exe hello_dumped.exe upx -d -q hello_dumped.exepython # sections.py — compare entry point and import count across two files import sys, pefile for p in sys.argv[1:]: pe = pefile.PE(p) ep = pe.OPTIONAL_HEADER.AddressOfEntryPoint named = [i.name.decode() for d in getattr(pe, "DIRECTORY_ENTRY_IMPORT", []) for i in d.imports if i.name] print(f"{p:16} entry_point={ep:#06x} imports={len(named)}")bash ./venv/bin/python sections.py hello.exe hello_upx.exe hello_dumped.exetext hello.exe entry_point=0x1440 imports=42 hello_upx.exe entry_point=0xc750 imports=12 hello_dumped.exe entry_point=0x1440 imports=42The repaired file's entry point now matches the OEP you found in step 2 exactly, and the import count is back to the original 42 — the same outcome Scylla's "Get Imports" + "Fix Dump" sequence produces by walking memory pointers rather than replaying a known algorithm. This is the validation step: a correct manual unpack should always reproduce this entry-point-and-import-count match.
-
Prove it behaves like the original, the final check from the lesson above:
bash wine hello_dumped.exetext hello, analystIdentical output to both the original and the packed version. For this lab that is a formality; for a real sample it is the difference between a usable dump and a plausible-looking one that crashes the moment you hand it to a disassembler.
Questions to answer: If hello_upx.exe had used a custom XOR-based
resolver instead of calling GetProcAddress for each import — the API-hashing
case — which of the numbers in step 3's table would still match, and which
would not? Why does the ESP-after-pushad trick still work even though this lab
found the OEP by static disassembly instead? If step 3's import count had come
back as 40 instead of 42, what would you check first in Scylla's invalid-import
list? Sketch, in one sentence each, what changes about this whole procedure if
the packer were a process-hollowing loader instead of a compressor.
Key takeaways
- Find the OEP with a hardware breakpoint on the stack address exposed right
after the stub's opening register save (ESP-after-pushad), or an execute
breakpoint on the region that just became writable-then-executable; for
multi-stage stubs, break on
VirtualAlloc/VirtualProtect/VirtualFreeand inspect each newly-executable region in turn. - A dump taken at the OEP is not runnable yet: in-memory section alignment does not match on-disk alignment, and the header's entry point field still points at the stub.
- Scylla and pe-sieve convert alignment and rewrite the entry point; Scylla's IAT search then finds pointer-shaped memory, resolves each pointer to a named export by address, and writes a fresh import directory.
- IAT rebuilding resolves pointers correctly even under API hashing, but cannot recover names the stub never looked up by name — expect an IAT of correct, unnamed functions in that case.
- Validate a fixed dump by checking the entry point and import count against what you expect, then by running it and comparing behaviour to the original — not by assuming the dump worked because it opened in a disassembler.
- Virtualised stubs, stolen bytes and TLS-callback-driven logic defeat this procedure outright; recognise them from the section table and imports before you invest time chasing an OEP that either does not exist or is not enough.