Lesson 10.1 · Beyond the EXE· 45 min
Shellcode Analysis
How to recognise raw shellcode without headers, follow its get-PC and API-resolution idioms, and disassemble, emulate and extract IOCs from it safely.
Objectives
- Distinguish position-independent shellcode from a PE by the absence of headers, imports and fixed addresses
- Recognise where shellcode hides and the byte patterns and entropy that betray it
- Read the get-PC, PEB-walking and hash-based API-resolution idioms in disassembly
- Disassemble a raw blob with Capstone and emulate it safely with Unicorn to log its behaviour
- Extract indicators from decoded shellcode and record them in a report
Sooner or later a sample hands you a blob of raw code with no file around it: a byte array pasted into a PowerShell one-liner, a stream inside a Word document, a resource that decrypts to something that is clearly instructions but is not a PE, or a region of private, executable memory in a dump. This is shellcode — a chunk of machine code written to run anywhere, with no loader, no linker and no operating system doing the setup a normal program relies on.
Everything earlier in this path assumed a file format: PE headers to parse, an import table to read, sections with known addresses. Shellcode throws all of that away. This lesson is about reading code that has no scaffolding: how to know you are looking at shellcode, how it finds its own bearings and calls Windows APIs, and how to disassemble, emulate and extract indicators from it without running it on anything you care about. The focus is analysis — recognising and understanding a blob someone else wrote, not authoring one.
What makes shellcode different from a PE
A normal Windows program is a PE file: the loader maps its sections, applies relocations, resolves its imports into the import address table, and jumps to a fixed entry point. Shellcode gets none of that. Three properties follow, and each one is a recognition cue.
| Property | A PE has | Shellcode has |
|---|---|---|
| Headers | MZ/PE signatures, a section table, a declared entry point | Nothing — the first byte is the first instruction |
| Position | A preferred base address and relocations to fix up | Position independence: it must work at whatever address it lands |
| API access | An import table the loader fills in | No imports; it must find every function it needs at runtime |
No headers means there is no metadata to tell you where code starts or how long it is. You are handed bytes and an implicit "execution begins at offset 0".
Position independence is the defining constraint. Shellcode cannot hardcode the address of its own data, because it does not know where it will be copied. It must compute addresses relative to wherever it is executing — which produces the tell-tale idioms in the next section.
No import table means it cannot call LoadLibrary or WinHTTPOpen the easy
way, because nobody resolved those names for it. It has to locate the base of
kernel32.dll in memory and walk export tables by hand. That machinery, absent
from ordinary code, is one of the clearest signs you are reading shellcode.
Where it turns up and how to recognise it
Shellcode rarely arrives as a neat file. It is usually a payload embedded inside another artefact, and part of triage is spotting the blob in the first place:
- Encoded in scripts. A PowerShell, JavaScript or VBScript loader carries the
bytes as Base64, a hex string or a comma-separated integer array, decodes them,
allocates executable memory (
VirtualAlloc,.NETMarshal.Copy) and jumps in. PowerShell, JavaScript and VBScript Malware covers those wrappers; what they wrap is shellcode. - In document objects. Malicious documents drop shellcode inside an OLE
stream, an RTF object or an exploit payload, often behind a
NOP sled (a run of
0x90or equivalent no-ops that gives an imprecise jump room to land). - In resources and overlays. A dropper stores an encrypted stage as a PE resource or appended overlay; decrypted, some stages are shellcode rather than a PE.
- In memory. During dynamic analysis you find a private,
RWXregion that was written and then executed. There is no file — the injected code exists only in the process. Dumping such regions is the subject of the memory-dumping lesson later in this path.
Heuristics that flag a candidate blob:
| Signal | What it looks like |
|---|---|
| Get-PC idiom | E8 00 00 00 00 (call $+5) followed by a pop, or a call/pop pair |
| PEB access | 65 48 8B ... — a mov with the 0x65 gs prefix (x64), or 64 8B ... with fs (x86) |
| fs/gs offsets | Reads of gs:[0x60] (x64 PEB) or fs:[0x30] (x86 PEB) |
| NOP sled | A long run of 0x90, or other single-byte no-ops |
| Printable-hash constants | 32-bit immediates compared in a loop (API-hash lookups) |
| Entropy | Encoded stubs raise entropy; a decoder stub sits in front of a high-entropy body |
None of these is proof on its own. The way to confirm a candidate is to disassemble it as raw code (below) and see whether it forms a coherent instruction stream that does position-independent things. A blob that disassembles into a get-PC idiom, a decode loop and a PEB walk is shellcode; a blob that disassembles into garbage that never resynchronises is probably data.
Tip: Always record which architecture and bitness you assumed. The same bytes decode into completely different instructions as x86 versus x86-64. If a blob looks like nonsense one way, try the other before concluding it is not code.
How shellcode finds its bearings
The get-PC idiom
Because it cannot reference its own data by absolute address, shellcode first
asks "where am I?". The classic answer abuses call: a
call pushes the address of the following instruction onto
the stack, so a call to the very next instruction,
followed by a pop, lands that address in a register.
call next ; pushes address of `next` (the byte after the call)
next:
pop rsi ; rsi now holds the runtime address of `next`You will also see E8 00 00 00 00 written inline (a call with a zero
displacement, i.e. to the next instruction) immediately before a pop. Once the
shellcode has one known runtime address, it computes every other address as an
offset from it. Seeing a call/pop with no matching ret, used purely to
load an address, is a strong shellcode signature. The lab at the end builds and
emulates exactly this shape.
Walking the PEB to find modules
To call an API, shellcode needs the base address of the DLL that exports it,
starting with kernel32.dll. It finds it through the Process Environment Block
(PEB), which the OS places at a fixed offset from a
segment register: gs:[0x60] on x86-64,
fs:[0x30] on x86. From the PEB it follows Ldr to the list of loaded modules
(InLoadOrderModuleList or InMemoryOrderModuleList), a linked list of every
DLL already mapped into the process. Walking that list, it reads each module's
base address and name until it reaches the one it wants.
In disassembly this reads as a mov through gs/fs, then a chain of pointer
dereferences at small constant offsets (general-purpose
registers holding the walk),
often followed by a loop comparing module names. You do not need to memorise the
offsets; recognising the gs:[0x60]/fs:[0x30] read followed by list traversal
is enough to know "this is resolving modules by hand", which only shellcode and
reflective loaders do.
Hash-based API resolution
Having found kernel32.dll, shellcode locates individual functions by parsing
the DLL's export directory in memory. To avoid carrying readable strings like
"WinExec" — which would show up in strings and in signatures — most modern
shellcode compares a hash of each export name against a precomputed constant.
This is API hashing, and it is worth reading that
technique page in full, because it is the single most common obstacle between you
and a shellcode's intent.
For the analyst, the pattern in disassembly is: a loop over the export name array, a small hashing routine (rotate-and-add or a similar mix), and a compare against a 32-bit immediate. Each such immediate corresponds to one API. You resolve them by running the same hash over a dictionary of Windows export names — or by letting an emulator do the resolution for you, which is usually faster.
Decoder stubs
Shellcode frequently carries its real body encoded, with a small decoder stub in front that rewrites the body in place before jumping to it. The stub uses the get-PC idiom to find the encoded bytes, loops over them applying an XOR, add or RC4 keystream, and then falls through into the decoded code. This is self-modifying code in miniature, and the XOR loop is the same construct you met in XOR string encryption, applied to instructions rather than strings. A static disassembler shows you only the stub; everything after it is high-entropy noise until the stub runs. That is precisely why emulation earns its place here.
Tools and workflow
Shellcode analysis alternates between static disassembly and controlled execution, exactly as Linear vs Recursive Disassembly described, but with no symbols and no headers to seed from.
Disassemble the raw blob. Load the bytes at a chosen base and decode from offset 0:
- Capstone in a short Python script is the fastest way to get a listing and to script recognition of the idioms above.
- Ghidra or IDA can import a flat binary: choose a raw/binary load, set the architecture (x86-64 or x86) and a base address, then disassemble from offset 0. You lose auto-analysis's reliance on headers, so you often mark the start manually and let recursive descent follow from there.
Emulate it safely. Because decoder stubs and hash resolution only reveal themselves at runtime, emulation is the workhorse — and it never executes the payload on your host:
- Unicorn runs the raw instructions in an isolated CPU with memory you map yourself. You hook execution, reads and writes, so you can watch a decoder stub rewrite its body and dump the result. Unicorn emulates the CPU only, not Windows, so API calls need stubbing; the automation module's emulation lesson builds that out.
- speakeasy (Mandiant) emulates the Windows environment as well as the CPU: it provides the PEB, fake DLLs and API implementations, so it can run shellcode through its module walk and hash resolution and log the APIs it would call, with arguments. For a real sample this is often the single most valuable step — point it at the blob, choose the architecture, and read the API trace.
- scdbg (built on libemu) is a long-standing, purpose-built shellcode emulator that reports the Win32 calls a blob makes. It is quick for a first pass on 32-bit shellcode and needs no scripting.
Debug it in the lab. When emulation stalls — an unusual API, an anti-analysis check, or code that reads state an emulator does not model — run the blob under a real debugger inside your isolated VM. A loader such as BlobRunner (OALabs) allocates executable memory, copies the blob in, prints the address and waits, so you can attach x64dbg, set a breakpoint on the entry and single-step. Do this only in the lab: BlobRunner really executes the shellcode.
Extract IOCs. Once decoded and emulated, harvest the indicators: URLs, C2
hosts and ports, user-agent strings, file paths, mutex names, and the set of APIs
it resolved (the hashes, mapped back to names). The API set alone is a capability
summary — VirtualAlloc + WinHttpConnect + CreateThread tells the story of a
downloader before you read a line of the payload.
Lab: decode a string from a raw x86-64 blob
You will build a tiny, benign, position-independent routine whose only job is to
XOR-decode an embedded URL-shaped string in place — a decoder-stub shape with no
API calls — then disassemble it with Capstone and emulate it with Unicorn,
hooking memory writes to watch the plaintext appear. The outputs below are real,
from Capstone 5.0.7 and Unicorn 2.1.4 in a Python 3.14 virtual environment on
macOS. Nothing is installed system-wide and nothing malicious runs: the "payload"
is a decode loop that stops at its own ret.
-
Create a virtual environment so the libraries stay local:
bash python3 -m venv venv ./venv/bin/pip install capstone unicornThe
keystone-engineassembler is the natural way to turn assembly into bytes, but its Python binding does not build against Python 3.14, so this lab hand-assembles the stub and verifies every byte with Capstone instead — which is exactly the discipline you want when you cannot trust an assembler. -
Save
build_shellcode.py. It lays out the stub by hand, XOR-encodes the string with key0x5A, writesstub.bin, and disassembles the code portion:python # build_shellcode.py - assemble a benign PIC XOR-decoder stub by hand, # then verify the bytes with Capstone. from capstone import Cs, CS_ARCH_X86, CS_MODE_64 KEY = 0x5A PLAIN = b"http://update.example.com/lab" DATALEN = len(PLAIN) encoded = bytes(b ^ KEY for b in PLAIN) stub = bytes([ 0xEB, 0x11, # 0x00 jmp short call_dec 0x5E, # 0x02 decode: pop rsi (rsi = &data) 0x48, 0x31, 0xC9, # 0x03 xor rcx, rcx 0xB1, DATALEN, # 0x06 mov cl, DATALEN 0x80, 0x36, KEY, # 0x08 next: xor byte ptr [rsi], KEY 0x48, 0xFF, 0xC6, # 0x0B inc rsi 0xFE, 0xC9, # 0x0E dec cl 0x75, 0xF6, # 0x10 jnz next 0xC3, # 0x12 ret (emu stops here) 0xE8, 0xEA, 0xFF, 0xFF, 0xFF, # 0x13 call_dec: call decode ]) blob = stub + encoded open("stub.bin", "wb").write(blob) print(f"blob hex : {blob.hex()}") print("=== Capstone disassembly (code portion) ===") md = Cs(CS_ARCH_X86, CS_MODE_64) for insn in md.disasm(stub, 0x1000): print(f"0x{insn.address:04x}: {insn.bytes.hex(' '):<20} " f"{insn.mnemonic} {insn.op_str}")The layout is a get-PC decoder stub: the opening
jmpskips forward to acallthat targetsdecode; thecallpushes the address of the byte after it — the start of the encoded data — whichpop rsiloads. The loop then XORs each byte in place. -
Run it:
bash ./venv/bin/python build_shellcode.pytext blob hex : eb115e4831c9b11d80365a48ffc6fec975f6c3e8eaffffff322e2e2a60... === Capstone disassembly (code portion) === 0x1000: eb 11 jmp 0x1013 0x1002: 5e pop rsi 0x1003: 48 31 c9 xor rcx, rcx 0x1006: b1 1d mov cl, 0x1d 0x1008: 80 36 5a xor byte ptr [rsi], 0x5a 0x100b: 48 ff c6 inc rsi 0x100e: fe c9 dec cl 0x1010: 75 f6 jne 0x1008 0x1012: c3 ret 0x1013: e8 ea ff ff ff call 0x1002Read it as an analyst would a real blob: the
callat0x1013targets0x1002, and the byte after thatcall(0x1018) is wherepop rsiwill point — the encoded data. Themov cl, 0x1dsets the length (29 bytes). There are no imports, no headers and no absolute data addresses: everything is relative to the popped program counter. -
Save
emulate_shellcode.py. It maps the blob and a stack, hooks every memory write, and stops at theretso nothing past the decode loop runs:python # emulate_shellcode.py - run the stub under Unicorn, hook writes, dump result. from unicorn import Uc, UC_ARCH_X86, UC_MODE_64, UC_HOOK_MEM_WRITE from unicorn.x86_const import UC_X86_REG_RSP BASE, STACK, PAGE = 0x1000000, 0x2000000, 0x1000 blob = open("stub.bin", "rb").read() DATA_OFF, DATALEN = 0x18, 29 RET_ADDR = BASE + 0x12 # the `ret`; emulation stops here mu = Uc(UC_ARCH_X86, UC_MODE_64) mu.mem_map(BASE, PAGE); mu.mem_map(STACK, PAGE) mu.mem_write(BASE, blob) mu.reg_write(UC_X86_REG_RSP, STACK + PAGE // 2) def region(addr): if BASE <= addr < BASE + PAGE: return f"blob+0x{addr-BASE:02x}" if STACK <= addr < STACK + PAGE: return "stack" return f"0x{addr:x}" def on_write(uc, access, address, size, value, _): b = value & 0xFF ch = chr(b) if 32 <= b < 127 else "." print(f" write {region(address):<12} = 0x{b:02x} '{ch}'") mu.hook_add(UC_HOOK_MEM_WRITE, on_write) mu.emu_start(BASE, RET_ADDR) data = bytes(mu.mem_read(BASE + DATA_OFF, DATALEN)) print(f"\ndecoded (hex) : {data.hex()}") print(f"decoded (ascii) : {data.decode('latin1')}") -
Run it:
bash ./venv/bin/python emulate_shellcode.pytext === emulating (memory writes hooked) === write stack = 0x18 '.' write blob+0x18 = 0x68 'h' write blob+0x19 = 0x74 't' write blob+0x1a = 0x74 't' write blob+0x1b = 0x70 'p' write blob+0x1c = 0x3a ':' write blob+0x1d = 0x2f '/' ... write blob+0x33 = 0x61 'a' write blob+0x34 = 0x62 'b' decoded (hex) : 687474703a2f2f7570646174652e6578616d706c652e636f6d2f6c6162 decoded (ascii) : http://update.example.com/labThe very first write is to the stack, not the data: it is the
callpushing its return address (0x...018, low byte0x18) — the get-PC idiom caught in the act. Every write afterwards lands in the data region, one byte at a time, and the URL assembles itself in memory. Dumping the region afterward gives you the plaintext IOC,http://update.example.com/lab, which a purely static look at the encoded blob would never reveal. -
For a real sample, the same shape scales up but you would not run it on your host. Instead:
- Run it through speakeasy (
speakeasy -t blob.bin -a x64 -r) or scdbg for 32-bit blobs, and read the logged API calls and their arguments — the module walk and hash resolution happen inside the emulator, and the C2 URL usually appears as an argument to a networking API. - When emulation stalls, load the blob with BlobRunner inside your isolated VM, attach x64dbg to the printed address, and single-step the decoder and the API-resolution loop.
- Run it through speakeasy (
Questions to answer: In the disassembly, how do you know 0x1018 is data and
not code, given there is no header to tell you? Why does the first memory write go
to the stack rather than the string, and which instruction caused it? If you
changed the XOR key in build_shellcode.py but not the length byte, what would
the emulated output look like, and how would you recover the key from the
encoded blob alone? For a real blob, why is speakeasy's API log often more useful
than a perfect static disassembly of the same bytes?
What to report
Fold shellcode findings into the triage report structure you already use:
- Where it was found and how it was extracted: the carrier (script, document, resource, memory region), the offset and length, and the decode or dump steps — hash the extracted blob separately, and label it a derived artefact.
- Architecture and shape: x86 or x86-64; whether it uses a get-PC idiom, a decoder stub (with the algorithm and key), a PEB walk and API hashing.
- Resolved APIs: the hash constants mapped back to function names, which give the capability summary.
- IOCs: URLs, hosts, ports, user-agents, mutexes, paths and file names recovered after decoding, each with how you obtained it so another analyst can reproduce it.
- Next stage: if the shellcode downloads or unpacks a further payload, note where it comes from and hash it once you have it.
Key takeaways
- Shellcode is position-independent code with no headers, no relocations and no import table; those three absences are how you recognise it against a PE.
- It shows up encoded in scripts, inside document objects, in resources and
overlays, and as executable memory in dumps; get-PC idioms,
gs:[0x60]/fs:[0x30]PEB reads, NOP sleds and hash constants flag a candidate blob. - It finds its bearings with the
call/popget-PC idiom, locates DLLs by walking the PEB module list, and resolves APIs by hash to avoid readable strings — read the API hashing page for that step. - Disassemble raw blobs with Capstone or a flat load in Ghidra/IDA, emulate them safely with Unicorn, speakeasy or scdbg to reveal decoder stubs and log API calls, and only debug live under BlobRunner inside an isolated VM.
- Report where the blob came from, its shape and algorithm, the APIs it resolves and the IOCs you decode, each with reproducible steps.