Skip to content

Lesson 9.1 · Evasion & Unpacking· 45 min

Anti-Disassembly

Recognising the tricks that make a disassembler lie — constant-condition jumps, overlapping instructions, rogue bytes — and repairing the listing.

Objectives

  • Explain which disassembler assumptions each anti-disassembly family attacks
  • Recognise constant-condition jumps, jump pairs, rogue bytes, overlapping instructions and obscured flow in Ghidra and IDA
  • Read the visible symptoms — red cross-references, undefined bytes, wrong function bounds — as anti-disassembly alarms
  • Repair a listing by redefining code and data, NOP-ing junk in the database and setting function bounds by hand
  • Write a small normaliser that patches a known anti-disassembly pattern in a copy of the sample

Linear vs Recursive Disassembly ended with a promise: the same weaknesses that make disassembly hard are the ones an author exploits on purpose. This lesson keeps that promise. Every trick here is an attack on an assumption your disassembler must make, and every trick has a visible symptom and a manual repair. The goal is not to admire the cleverness; it is to notice within seconds that a listing is fiction, name which family produced it, and get back to a correct control-flow graph.

Anti-disassembly is cheap for the author and expensive for you only if it surprises you. Once you recognise the handful of families below, most of it becomes a mechanical clean-up step before the real analysis begins.

Why anti-disassembly works

Recall the two things a static disassembler must infer on variable-length x86-64 code: where code is, and where each instruction begins. Recursive descent — the engine under Ghidra, IDA and Binary Ninja — answers both by following control flow from known starting points. That gives it three exploitable assumptions:

  1. Both edges of a conditional branch are reachable. The engine decodes the taken target and the fall-through, because statically it cannot prove either is dead.
  2. A call returns to the instruction after it, and a ret returns to its caller. Function bounds and cross-references are built on this.
  3. Cross-references come from direct branches. A target reached only through a register or a computed address leaves no visible xref.

Each family below turns one of these assumptions into a weapon. The x86 instruction encoding is dense — most byte values decode as something — so a single misplaced byte does not raise an error; it silently produces plausible nonsense.

The families

Jumps with a constant condition

The signature trick, and the one you meet first. The author zeroes a flag and then branches on it, so the conditional jump is unconditional in practice:

asm
xor  eax, eax        ; ZF = 1, always
jz   real            ; always taken
.byte 0xE8           ; rogue byte the CPU never reaches
real:
lea  eax, [rcx+1]    ; the real body
ret

The CPU always takes the branch, so the byte after the jz never executes. But recursive descent decodes the fall-through anyway, starting at 0xE8 — the opcode of a five-byte call rel32. That bogus call swallows the real lea/ret, and the branch target now lands inside it. This is an opaque predicate: a condition whose value is fixed but not obvious to a tool. xor/jz, or eax,eax/jnz, test esp,esp /jns, cmp against a constant the code just set — all are variations.

Two jumps to the same target

The same effect with two branches instead of a fake flag: jz L immediately followed by jnz L. Between them they cover both flag states, so L is always reached, yet the engine treats the jnz fall-through as live and decodes the rogue byte sitting there. The disassembler never combines two instructions into one unconditional jump, so it cannot see that the fall-through is dead.

Junk bytes after unconditional transfers

After a jmp, a ret, or a non-returning call (to ExitProcess, abort), nothing should be decoded as fall-through — but recursive descent's heuristics, and every linear sweep, will decode whatever bytes follow. Authors place a multi-byte opcode there (0xE8, 0xE9, 0x0F) so the "instruction" runs into and consumes the real code that follows. The defence a normal binary gives you — .pdata function bounds on x64 Windows (see PE Sections and Functions and Control-Flow Graphs) — is exactly what hand-written junk lacks.

Overlapping instructions

The hardest family to represent. A single byte is deliberately the last byte of one instruction and the first byte of the next one actually executed. The classic seed is EB FF: a two-byte jmp whose target is its own second byte, where FF C0 decodes as inc eax. No flat listing can show one byte as part of two instructions, so any linear representation must hide one of the two real instructions. This is impossible disassembly in the strict sense — the fix is not a cleaner listing but a note, or replacing the whole sequence with nops in your database once you understand its net effect.

Return-address and exception-based control flow

call/ret are not just for functions. A call $+5 pushes the address of the next byte; the code then adds a constant to [rsp] and executes ret, jumping to a computed location with no xref and no caller relationship. The engine ends the function early at the ret and never connects it to the real continuation. Structured Exception Handling is the same idea escalated: the sample registers a handler, triggers a fault (a divide-by-zero, an int3), and resumes inside the handler — control flow the disassembler cannot follow at all, because the transfer happens through the OS. The handler body appears as unreferenced data.

Function-pointer indirection

Storing a function's address in a variable and calling through it (call [rbp-8], call rax) is ordinary C, and it hides the edge. The engine records the one place the address was taken but not the places it is called, so the callee shows too few cross-references and its prototype never propagates. Combined with control-flow flattening, where a dispatcher switch drives every basic block through an indirect jump, this turns a readable function into a soup of blocks with no visible order.

Stack-frame confusion

Advanced disassemblers reconstruct a function's locals and arguments by tracking rsp/rbp adjustments. Deliberately unbalanced or misleading stack arithmetic — allocating with a value that depends on a branch, referencing [rsp+N] with an N chosen so the tool computes the wrong frame — defeats that analysis. The symptom is a function whose decompilation is full of undefined variables and whose stack pointer the tool reports as "unbalanced"; it also disables the decompiler for that function.

Reading the symptoms

You rarely need to know which family you are facing before you know that something is wrong. Each one leaves a visible mark:

FamilyAssumption attackedWhat you see in Ghidra / IDA
Constant-condition jumpBoth branch edges liveBranch target drawn in red (points inside an instruction); a call/jmp to a wild address
Two jumps, one targetBoth edges liveSame red xref; two branches stacked on the same line region
Junk after jmp/retFall-through is codeA too-long instruction straddling a function end; undefined bytes after it
Overlapping instructionsOne byte, one instructionA branch target with no instruction start; bytes claimed by two decodings
ret/SEH flow controlcall returns, ret returns to callerFunction ends early; a real routine shown as unreferenced data; "sp-analysis failed"
Function-pointer indirectionXrefs come from direct branchesCallee with too few xrefs; no argument types propagated
Stack-frame confusionBalanced, analysable frameundefined locals; unbalanced-stack warning; decompiler refuses the function

Tip: A branch target rendered in red in IDA, or an "Instruction does not fall through" / bad-instruction bookmark in Ghidra, is the single fastest anti-disassembly alarm. Jump to that address, look at the bytes, and you will almost always find one of the families above.

The analyst's fixes

There is a strict order of preference, from least to most invasive. Crucially, you edit your analysis database, not the sample on disk — the file must stay byte-for-byte identical for hashing, YARA and later dynamic runs.

  1. Redefine code and data. In IDA, U undefines, C turns bytes into code, D turns them into data. In Ghidra, Clear Code Bytes then Disassemble (D), or Data → choose type. Turn the rogue byte into data so the engine resumes decoding at the real target. This alone fixes the first three families.
  2. NOP the junk in the database. For overlapping or multi-level sequences where redefining is not enough, patch the confusing bytes to 0x90 in the tool (IDA: Edit → Patch program → Change byte; Ghidra: Patch Instruction). Record the net effect first — an EB FF C0 48 sequence is just inc eax then dec eax, an elaborate no-op.
  3. Set function bounds by hand. When a rogue ret ends a function early, or junk merges two functions, define the real bounds: IDA Alt+P (Edit function) or create a function with P; Ghidra Create Function / Edit Function. Add the missing cross-reference so the callee reconnects to the graph.
  4. Script the repair for repeated patterns. Obfuscators apply the same pattern hundreds of times. Write a short script that finds it and normalises every instance — an IDAPython / Ghidra script that patches bytes and re-disassembles, or a standalone Capstone
    • pefile pass over a copy of the file. The lab builds one.

Warning: Patching the database changes your view, not the program. If you then run the sample in a debugger, its bytes are the original ones — your NOPs are not there. Keep the two models separate in your notes.

Lab: two disassemblers, one rogue byte

You need x86_64-w64-mingw32-gcc (or clang -target x86_64-w64-windows-gnu), its objdump, Python 3 with a virtual environment holding capstone and pefile, and optionally Wine to run the result. The program is harmless: it prints one number. Listings below come from MinGW-w64 GCC 15.2.0, binutils objdump, Capstone 5.0.7 and pefile 2024.8.26; addresses differ with other toolchains.

  1. Save adlab.c. The trick is written in a top-level asm block so the exact bytes survive the compiler:

    c
    /* adlab.c: a harmless function hidden behind a constant-condition jump.
       guarded(x) returns x + 1; the jz is always taken, so the 0xE8 rogue
       byte after it never executes. Windows x64 passes the first int in ecx. */
    #include <stdio.h>
    
    extern int guarded(int x);
    
    __asm__(
        ".intel_syntax noprefix\n"
        ".text\n"
        ".globl guarded\n"
    "guarded:\n"
        "    xor   eax, eax\n"   /* ZF = 1, unconditionally */
        "    jz    1f\n"         /* always taken: skip the rogue byte     */
        "    .byte 0xE8\n"       /* rogue byte: opcode of call rel32       */
    "1:  lea   eax, [rcx+1]\n"   /* the real body the trick hides          */
        "    ret\n"
        ".att_syntax prefix\n"
    );
    
    int main(void) {
        printf("guarded(41) = %d\n", guarded(41));
        return 0;
    }
    bash
    x86_64-w64-mingw32-gcc -O2 -o adlab.exe adlab.c
    python3 -m venv venv && ./venv/bin/pip install capstone pefile
    wine adlab.exe          # optional: prints  guarded(41) = 42
  2. Linear sweep desynchronises. Disassemble with objdump and read guarded:

    bash
    x86_64-w64-mingw32-objdump -d -M intel adlab.exe | grep -A 6 '<guarded>:'
    text
    0000000140001490 <guarded>:
       140001490:	31 c0                	xor    eax,eax
       140001492:	74 01                	je     140001495 <guarded+0x5>
       140001494:	e8 8d 41 01 c3       	call   103015626 <__size_of_stack_reserve__+0x102e15626>
       140001499:	90                   	nop
       14000149a:	90                   	nop

    The je at ...492 points to ...495, which is inside the five-byte call at ...494. That mid-instruction target is the alarm. The real body — lea eax,[rcx+1]; ret — is gone, its bytes (8d 41 01 c3) absorbed into the bogus call to a wild address.

  3. Recursive descent follows the jump. Save recdis_pe.py, a minimal recursive walker that decodes one function, follows direct branch targets, and flags any two instructions that claim the same byte:

    python
    # recdis_pe.py: tiny recursive-descent disassembler for one function in a PE.
    # usage: python recdis_pe.py file.exe <rva-hex>
    import sys
    import pefile
    from capstone import Cs, CS_ARCH_X86, CS_MODE_64, CS_GRP_JUMP, CS_GRP_CALL, CS_GRP_RET
    from capstone.x86 import X86_OP_IMM
    
    STOP = {"jmp", "ret", "hlt"}
    
    pe = pefile.PE(sys.argv[1])
    base = pe.OPTIONAL_HEADER.ImageBase
    start = base + int(sys.argv[2], 16)
    image = pe.get_memory_mapped_image()
    lo, hi = start, start + 0x20              # small window around the function
    
    md = Cs(CS_ARCH_X86, CS_MODE_64)
    md.detail = True
    
    def target(insn):
        op = insn.operands[0] if insn.operands else None
        return op.imm if op and op.type == X86_OP_IMM else None
    
    insns, todo = {}, [start]
    while todo:
        addr = todo.pop()
        while lo <= addr < hi and addr not in insns:
            i = next(md.disasm(image[addr - base:], addr, 1), None)
            if i is None:
                break
            insns[addr] = i
            if i.group(CS_GRP_JUMP) or i.group(CS_GRP_CALL):
                t = target(i)
                if t is not None:
                    todo.append(t)
            if i.mnemonic in STOP or i.group(CS_GRP_RET):
                break
            addr += i.size
    
    covered = {}
    for a, i in insns.items():
        for b in range(a, a + i.size):
            covered.setdefault(b, []).append(a)
    for a in sorted(insns):
        i = insns[a]
        clash = {o for b in range(a, a + i.size) for o in covered[b]} - {a}
        flag = "   <-- overlaps " + ", ".join(hex(c) for c in sorted(clash)) if clash else ""
        print(f"{a:#x}: {i.bytes.hex(' '):<17} {i.mnemonic} {i.op_str}{flag}")
    bash
    ./venv/bin/python recdis_pe.py adlab.exe 0x1490
    text
    0x140001490: 31 c0             xor eax, eax
    0x140001492: 74 01             je 0x140001495
    0x140001494: e8 8d 41 01 c3    call 0x103015626   <-- overlaps 0x140001495, 0x140001498
    0x140001495: 8d 41 01          lea eax, [rcx + 1]   <-- overlaps 0x140001494
    0x140001498: c3                ret    <-- overlaps 0x140001494

    By following the je target, the walker decodes the real lea eax,[rcx+1]; ret that linear sweep lost, and marks the overlap with the bogus call. This is what Ghidra and IDA do internally — and why the conflict shows up in red rather than being silently swallowed.

  4. Normalise the pattern. Redefining bytes by hand is fine once; for a whole sample, script it. Save normalise.py, which finds the xor eax,eax; jz +1 signature and NOPs the byte the jump skips, in a copy of the file:

    python
    # normalise.py: neutralise the "constant-condition jz over a rogue byte" pattern.
    # Finds  xor eax,eax ; jz +1  (31 C0 74 01)  and overwrites the byte the jump
    # skips with a NOP (0x90), in a copy of the file. usage: python normalise.py in.exe out.exe
    import sys
    
    PATTERN = bytes.fromhex("31c07401")     # xor eax,eax ; jz .+3 (skip 1 byte)
    data = bytearray(open(sys.argv[1], "rb").read())
    
    patched, off = 0, 0
    while True:
        off = data.find(PATTERN, off)
        if off == -1:
            break
        rogue = off + len(PATTERN)           # the single byte the jz skips
        if data[rogue] != 0x90:
            print(f"file offset {rogue:#x}: {data[rogue]:#04x} -> 0x90")
            data[rogue] = 0x90
            patched += 1
        off = rogue
    
    open(sys.argv[2], "wb").write(data)
    print(f"{patched} rogue byte(s) patched")
    bash
    ./venv/bin/python normalise.py adlab.exe adlab_clean.exe
    text
    file offset 0xa94: 0xe8 -> 0x90
    1 rogue byte(s) patched
  5. Read the clean copy. Linear sweep now resynchronises immediately, because the rogue 0xE8 is a one-byte nop and the je target lands on a real instruction:

    bash
    x86_64-w64-mingw32-objdump -d -M intel adlab_clean.exe | grep -A 5 '<guarded>:'
    text
    0000000140001490 <guarded>:
       140001490:	31 c0                	xor    eax,eax
       140001492:	74 01                	je     140001495 <guarded+0x5>
       140001494:	90                   	nop
       140001495:	8d 41 01             	lea    eax,[rcx+0x1]
       140001498:	c3                   	ret

    The patched copy still runs (wine adlab_clean.exe prints guarded(41) = 42) because the rogue byte was dead code. That is the point: your normaliser changed a copy for readability, and left the original untouched for hashing and dynamic analysis.

Questions to answer: In step 2, why does objdump not print an error on the bogus call even though its target is a wild address? If the rogue byte had been 0x0F (the start of a two-byte opcode) instead of 0xE8, how many real bytes would the linear listing have swallowed? The normaliser matches only the exact 31 C0 74 01 pattern — name two constant-condition variants it would miss, and how you would generalise it. Why is patching a copy on disk the wrong choice for a sample you still need to run in a debugger, and what should you patch instead?

Key takeaways

  • Every anti-disassembly family attacks a specific disassembler assumption: both branch edges live, call/ret behave normally, or xrefs come from direct branches.
  • Constant-condition jumps, jump pairs and rogue bytes hide real code behind a fake multi-byte instruction; a branch target that lands inside an instruction (red in IDA) is the fastest alarm.
  • Overlapping instructions and ret/SEH flow control cannot always be shown in a flat listing; recognise the net effect and note it rather than chasing a perfect view.
  • Fix the database, not the file: redefine code and data, NOP junk in the tool, set function bounds and add missing cross-references by hand.
  • For patterns applied at scale, script the normalisation over a copy — and keep the original bytes intact for hashing and dynamic analysis.