Leçon 9.1 · Évasion & unpacking· 45 min
Anti-Disassembly
Recognising the tricks that make a disassembler lie — constant-condition jumps, overlapping instructions, rogue bytes — and repairing the listing.
Cette leçon n’est disponible qu’en anglais pour le moment.
Objectifs
- Explain which disassembler assumptions each anti-disassembly family attacks
- Recognise constant-condition jumps, jump pairs, rogue bytes, overlapping instructions and obscured flow in Ghidra and IDA
- Read the visible symptoms — red cross-references, undefined bytes, wrong function bounds — as anti-disassembly alarms
- Repair a listing by redefining code and data, NOP-ing junk in the database and setting function bounds by hand
- Write a small normaliser that patches a known anti-disassembly pattern in a copy of the sample
Linear vs Recursive Disassembly ended with a promise: the same weaknesses that make disassembly hard are the ones an author exploits on purpose. This lesson keeps that promise. Every trick here is an attack on an assumption your disassembler must make, and every trick has a visible symptom and a manual repair. The goal is not to admire the cleverness; it is to notice within seconds that a listing is fiction, name which family produced it, and get back to a correct control-flow graph.
Anti-disassembly is cheap for the author and expensive for you only if it surprises you. Once you recognise the handful of families below, most of it becomes a mechanical clean-up step before the real analysis begins.
Why anti-disassembly works
Recall the two things a static disassembler must infer on variable-length x86-64 code: where code is, and where each instruction begins. Recursive descent — the engine under Ghidra, IDA and Binary Ninja — answers both by following control flow from known starting points. That gives it three exploitable assumptions:
- Both edges of a conditional branch are reachable. The engine decodes the taken target and the fall-through, because statically it cannot prove either is dead.
- A
callreturns to the instruction after it, and aretreturns to its caller. Function bounds and cross-references are built on this. - Cross-references come from direct branches. A target reached only through a register or a computed address leaves no visible xref.
Each family below turns one of these assumptions into a weapon. The x86 instruction encoding is dense — most byte values decode as something — so a single misplaced byte does not raise an error; it silently produces plausible nonsense.
The families
Jumps with a constant condition
The signature trick, and the one you meet first. The author zeroes a flag and then branches on it, so the conditional jump is unconditional in practice:
xor eax, eax ; ZF = 1, always
jz real ; always taken
.byte 0xE8 ; rogue byte the CPU never reaches
real:
lea eax, [rcx+1] ; the real body
retThe CPU always takes the branch, so the byte after the jz never executes. But
recursive descent decodes the fall-through anyway, starting at 0xE8 — the
opcode of a five-byte call rel32. That bogus call swallows the real
lea/ret, and the branch target now lands inside it. This is an
opaque predicate: a condition whose value is
fixed but not obvious to a tool. xor/jz, or eax,eax/jnz, test esp,esp
/jns, cmp against a constant the code just set — all are variations.
Two jumps to the same target
The same effect with two branches instead of a fake flag: jz L immediately
followed by jnz L. Between them they cover both flag states, so L is always
reached, yet the engine treats the jnz fall-through as live and decodes the
rogue byte sitting there. The disassembler never combines two instructions into
one unconditional jump, so it cannot see that the fall-through is dead.
Junk bytes after unconditional transfers
After a jmp, a ret, or a non-returning call (to ExitProcess, abort),
nothing should be decoded as fall-through — but recursive descent's heuristics,
and every linear sweep, will decode whatever bytes follow. Authors place a
multi-byte opcode there (0xE8, 0xE9, 0x0F) so the "instruction" runs into
and consumes the real code that follows. The defence a normal binary gives you —
.pdata function bounds on x64 Windows (see PE Sections
and Functions and Control-Flow Graphs) — is
exactly what hand-written junk lacks.
Overlapping instructions
The hardest family to represent. A single byte is deliberately the last byte of
one instruction and the first byte of the next one actually executed. The
classic seed is EB FF: a two-byte jmp whose target is its own second byte,
where FF C0 decodes as inc eax. No flat listing can show one byte as part of
two instructions, so any linear representation must hide one of the two real
instructions. This is impossible disassembly in the strict sense — the fix is
not a cleaner listing but a note, or replacing the whole sequence with
nops in your database once you understand its net effect.
Return-address and exception-based control flow
call/ret are not just for functions. A call $+5 pushes the address of the
next byte; the code then adds a constant to [rsp] and executes ret, jumping
to a computed location with no xref and no caller relationship. The engine
ends the function early at the ret and never connects it to the real
continuation. Structured Exception Handling is the same idea escalated: the
sample registers a handler, triggers a fault (a divide-by-zero, an int3), and
resumes inside the handler — control flow the disassembler cannot follow at all,
because the transfer happens through the OS. The handler body appears as
unreferenced data.
Function-pointer indirection
Storing a function's address in a variable and calling through it (call [rbp-8],
call rax) is ordinary C, and it hides the edge. The engine records the one
place the address was taken but not the places it is called, so the callee
shows too few cross-references and its prototype never propagates. Combined with
control-flow flattening, where a
dispatcher switch drives every basic block through an indirect jump, this
turns a readable function into a soup of blocks with no visible order.
Stack-frame confusion
Advanced disassemblers reconstruct a function's locals and arguments by tracking
rsp/rbp adjustments. Deliberately unbalanced or misleading stack arithmetic —
allocating with a value that depends on a branch, referencing [rsp+N] with an
N chosen so the tool computes the wrong frame — defeats that analysis. The
symptom is a function whose decompilation is full of undefined variables and
whose stack pointer the tool reports as "unbalanced"; it also disables the
decompiler for that function.
Reading the symptoms
You rarely need to know which family you are facing before you know that something is wrong. Each one leaves a visible mark:
| Family | Assumption attacked | What you see in Ghidra / IDA |
|---|---|---|
| Constant-condition jump | Both branch edges live | Branch target drawn in red (points inside an instruction); a call/jmp to a wild address |
| Two jumps, one target | Both edges live | Same red xref; two branches stacked on the same line region |
Junk after jmp/ret | Fall-through is code | A too-long instruction straddling a function end; undefined bytes after it |
| Overlapping instructions | One byte, one instruction | A branch target with no instruction start; bytes claimed by two decodings |
ret/SEH flow control | call returns, ret returns to caller | Function ends early; a real routine shown as unreferenced data; "sp-analysis failed" |
| Function-pointer indirection | Xrefs come from direct branches | Callee with too few xrefs; no argument types propagated |
| Stack-frame confusion | Balanced, analysable frame | undefined locals; unbalanced-stack warning; decompiler refuses the function |
Tip: A branch target rendered in red in IDA, or an "Instruction does not fall through" / bad-instruction bookmark in Ghidra, is the single fastest anti-disassembly alarm. Jump to that address, look at the bytes, and you will almost always find one of the families above.
The analyst's fixes
There is a strict order of preference, from least to most invasive. Crucially, you edit your analysis database, not the sample on disk — the file must stay byte-for-byte identical for hashing, YARA and later dynamic runs.
- Redefine code and data. In IDA,
Uundefines,Cturns bytes into code,Dturns them into data. In Ghidra, Clear Code Bytes then Disassemble (D), or Data → choose type. Turn the rogue byte into data so the engine resumes decoding at the real target. This alone fixes the first three families. - NOP the junk in the database. For overlapping or multi-level sequences
where redefining is not enough, patch the confusing bytes to
0x90in the tool (IDA: Edit → Patch program → Change byte; Ghidra: Patch Instruction). Record the net effect first — anEB FF C0 48sequence is justinc eaxthendec eax, an elaborate no-op. - Set function bounds by hand. When a rogue
retends a function early, or junk merges two functions, define the real bounds: IDAAlt+P(Edit function) or create a function withP; Ghidra Create Function / Edit Function. Add the missing cross-reference so the callee reconnects to the graph. - Script the repair for repeated patterns. Obfuscators apply the same
pattern hundreds of times. Write a short script that finds it and normalises
every instance — an IDAPython / Ghidra script that patches bytes and
re-disassembles, or a standalone Capstone
- pefile pass over a copy of the file. The lab builds one.
Warning: Patching the database changes your view, not the program. If you then run the sample in a debugger, its bytes are the original ones — your NOPs are not there. Keep the two models separate in your notes.
Lab: two disassemblers, one rogue byte
You need x86_64-w64-mingw32-gcc (or clang -target x86_64-w64-windows-gnu),
its objdump, Python 3 with a virtual environment holding capstone and
pefile, and optionally Wine to run the result. The program is harmless: it
prints one number. Listings below come from MinGW-w64 GCC 15.2.0, binutils
objdump, Capstone 5.0.7 and pefile 2024.8.26; addresses differ with other
toolchains.
-
Save
adlab.c. The trick is written in a top-levelasmblock so the exact bytes survive the compiler:c /* adlab.c: a harmless function hidden behind a constant-condition jump. guarded(x) returns x + 1; the jz is always taken, so the 0xE8 rogue byte after it never executes. Windows x64 passes the first int in ecx. */ #include <stdio.h> extern int guarded(int x); __asm__( ".intel_syntax noprefix\n" ".text\n" ".globl guarded\n" "guarded:\n" " xor eax, eax\n" /* ZF = 1, unconditionally */ " jz 1f\n" /* always taken: skip the rogue byte */ " .byte 0xE8\n" /* rogue byte: opcode of call rel32 */ "1: lea eax, [rcx+1]\n" /* the real body the trick hides */ " ret\n" ".att_syntax prefix\n" ); int main(void) { printf("guarded(41) = %d\n", guarded(41)); return 0; }bash x86_64-w64-mingw32-gcc -O2 -o adlab.exe adlab.c python3 -m venv venv && ./venv/bin/pip install capstone pefile wine adlab.exe # optional: prints guarded(41) = 42 -
Linear sweep desynchronises. Disassemble with
objdumpand readguarded:bash x86_64-w64-mingw32-objdump -d -M intel adlab.exe | grep -A 6 '<guarded>:'text 0000000140001490 <guarded>: 140001490: 31 c0 xor eax,eax 140001492: 74 01 je 140001495 <guarded+0x5> 140001494: e8 8d 41 01 c3 call 103015626 <__size_of_stack_reserve__+0x102e15626> 140001499: 90 nop 14000149a: 90 nopThe
jeat...492points to...495, which is inside the five-bytecallat...494. That mid-instruction target is the alarm. The real body —lea eax,[rcx+1]; ret— is gone, its bytes (8d 41 01 c3) absorbed into the boguscallto a wild address. -
Recursive descent follows the jump. Save
recdis_pe.py, a minimal recursive walker that decodes one function, follows direct branch targets, and flags any two instructions that claim the same byte:python # recdis_pe.py: tiny recursive-descent disassembler for one function in a PE. # usage: python recdis_pe.py file.exe <rva-hex> import sys import pefile from capstone import Cs, CS_ARCH_X86, CS_MODE_64, CS_GRP_JUMP, CS_GRP_CALL, CS_GRP_RET from capstone.x86 import X86_OP_IMM STOP = {"jmp", "ret", "hlt"} pe = pefile.PE(sys.argv[1]) base = pe.OPTIONAL_HEADER.ImageBase start = base + int(sys.argv[2], 16) image = pe.get_memory_mapped_image() lo, hi = start, start + 0x20 # small window around the function md = Cs(CS_ARCH_X86, CS_MODE_64) md.detail = True def target(insn): op = insn.operands[0] if insn.operands else None return op.imm if op and op.type == X86_OP_IMM else None insns, todo = {}, [start] while todo: addr = todo.pop() while lo <= addr < hi and addr not in insns: i = next(md.disasm(image[addr - base:], addr, 1), None) if i is None: break insns[addr] = i if i.group(CS_GRP_JUMP) or i.group(CS_GRP_CALL): t = target(i) if t is not None: todo.append(t) if i.mnemonic in STOP or i.group(CS_GRP_RET): break addr += i.size covered = {} for a, i in insns.items(): for b in range(a, a + i.size): covered.setdefault(b, []).append(a) for a in sorted(insns): i = insns[a] clash = {o for b in range(a, a + i.size) for o in covered[b]} - {a} flag = " <-- overlaps " + ", ".join(hex(c) for c in sorted(clash)) if clash else "" print(f"{a:#x}: {i.bytes.hex(' '):<17} {i.mnemonic} {i.op_str}{flag}")bash ./venv/bin/python recdis_pe.py adlab.exe 0x1490text 0x140001490: 31 c0 xor eax, eax 0x140001492: 74 01 je 0x140001495 0x140001494: e8 8d 41 01 c3 call 0x103015626 <-- overlaps 0x140001495, 0x140001498 0x140001495: 8d 41 01 lea eax, [rcx + 1] <-- overlaps 0x140001494 0x140001498: c3 ret <-- overlaps 0x140001494By following the
jetarget, the walker decodes the reallea eax,[rcx+1]; retthat linear sweep lost, and marks the overlap with the boguscall. This is what Ghidra and IDA do internally — and why the conflict shows up in red rather than being silently swallowed. -
Normalise the pattern. Redefining bytes by hand is fine once; for a whole sample, script it. Save
normalise.py, which finds thexor eax,eax; jz +1signature and NOPs the byte the jump skips, in a copy of the file:python # normalise.py: neutralise the "constant-condition jz over a rogue byte" pattern. # Finds xor eax,eax ; jz +1 (31 C0 74 01) and overwrites the byte the jump # skips with a NOP (0x90), in a copy of the file. usage: python normalise.py in.exe out.exe import sys PATTERN = bytes.fromhex("31c07401") # xor eax,eax ; jz .+3 (skip 1 byte) data = bytearray(open(sys.argv[1], "rb").read()) patched, off = 0, 0 while True: off = data.find(PATTERN, off) if off == -1: break rogue = off + len(PATTERN) # the single byte the jz skips if data[rogue] != 0x90: print(f"file offset {rogue:#x}: {data[rogue]:#04x} -> 0x90") data[rogue] = 0x90 patched += 1 off = rogue open(sys.argv[2], "wb").write(data) print(f"{patched} rogue byte(s) patched")bash ./venv/bin/python normalise.py adlab.exe adlab_clean.exetext file offset 0xa94: 0xe8 -> 0x90 1 rogue byte(s) patched -
Read the clean copy. Linear sweep now resynchronises immediately, because the rogue
0xE8is a one-bytenopand thejetarget lands on a real instruction:bash x86_64-w64-mingw32-objdump -d -M intel adlab_clean.exe | grep -A 5 '<guarded>:'text 0000000140001490 <guarded>: 140001490: 31 c0 xor eax,eax 140001492: 74 01 je 140001495 <guarded+0x5> 140001494: 90 nop 140001495: 8d 41 01 lea eax,[rcx+0x1] 140001498: c3 retThe patched copy still runs (
wine adlab_clean.exeprintsguarded(41) = 42) because the rogue byte was dead code. That is the point: your normaliser changed a copy for readability, and left the original untouched for hashing and dynamic analysis.
Questions to answer: In step 2, why does objdump not print an error on the
bogus call even though its target is a wild address? If the rogue byte had been
0x0F (the start of a two-byte opcode) instead of 0xE8, how many real bytes
would the linear listing have swallowed? The normaliser matches only the exact
31 C0 74 01 pattern — name two constant-condition variants it would miss, and
how you would generalise it. Why is patching a copy on disk the wrong choice for
a sample you still need to run in a debugger, and what should you patch instead?
Key takeaways
- Every anti-disassembly family attacks a specific disassembler assumption:
both branch edges live,
call/retbehave normally, or xrefs come from direct branches. - Constant-condition jumps, jump pairs and rogue bytes hide real code behind a fake multi-byte instruction; a branch target that lands inside an instruction (red in IDA) is the fastest alarm.
- Overlapping instructions and
ret/SEH flow control cannot always be shown in a flat listing; recognise the net effect and note it rather than chasing a perfect view. - Fix the database, not the file: redefine code and data, NOP junk in the tool, set function bounds and add missing cross-references by hand.
- For patterns applied at scale, script the normalisation over a copy — and keep the original bytes intact for hashing and dynamic analysis.