Leçon 3.2 · Triage statique· 45 min
Strings and Obfuscated Strings
Extract ASCII and UTF-16 strings, separate indicators from runtime noise, and recover stack, XOR and decoded strings with FLOSS and CyberChef.
Cette leçon n’est disponible qu’en anglais pour le moment.
Objectifs
- Extract ASCII and UTF-16LE strings with GNU strings, Sysinternals strings and FLOSS
- Separate analyst-relevant strings from compiler and runtime noise
- Recognise stack strings, single-byte XOR, Base64 and RC4 as ways of hiding strings
- Explain at a high level how FLOSS recovers stack, tight and decoded strings
- Brute-force single-byte XOR with a known-plaintext crib
After identity and hashes, the next triage step is reading the text the file carries. A C2 URL, a mutex name, a ransom note or a developer's PDB path can tell you more in thirty seconds than an hour in a disassembler. That is also why malware authors hide their strings, and why noticing missing strings is as important as reading the ones that are there.
You already used strings in earlier labs. This lesson makes it a method:
which encodings to search, how to cut the noise, what to look for, and what to
do when the interesting text is not in the file in readable form.
What a "string" is to a tool
A strings extractor knows nothing about the file format. It scans the raw bytes for runs of printable characters at least N long and prints each run. Two encodings matter on Windows:
"Lab" as ASCII / UTF-8 4C 61 62 00
"Lab" as UTF-16LE (wchar_t) 4C 00 61 00 62 00 00 00Windows APIs ending in W (CreateFileW, RegSetValueExW) take UTF-16LE,
and so do L"..." literals in C, .NET string tables, resources and the
version-info block. An ASCII scanner sees L, a zero byte, a, a zero byte
and gives up, because each run is only one character long. Always search
both encodings.
| Tool | ASCII | UTF-16LE | Notes |
|---|---|---|---|
GNU strings (binutils) | default | -e l | Default minimum length 4; -t x prints hex offsets |
Sysinternals strings | -a only | -u only | Searches both by default; minimum length 3; -o prints offsets |
| FLOSS | yes | yes | Also recovers stack, tight and decoded strings |
On macOS the built-in strings is Apple's and has no -e option. Use the one
from binutils or mingw-w64 (x86_64-w64-mingw32-strings).
Minimum length and noise
The minimum length is a noise dial. Machine code and compressed data contain
short printable runs by pure chance. On a small mingw-w64 test program, the
ASCII run count drops from about 200 at -n 4 to about 115 at -n 6 and under
100 at -n 8, and none of the lost strings were interesting. Start at 6, and
drop to 4 only when you are hunting something specific, such as short mutex
names or two-letter commands.
Random-looking short runs such as AWAVAUATUWVSH or [^_]A\A]A^A_ are
x86-64 instructions: the REX-prefixed push and pop of callee-saved
registers in function prologues and epilogues happen to be printable. Learn to
recognise them and skip them.
Triage: signal versus noise
Most strings in a binary were not written by the author. They come from the compiler, the C runtime and statically linked libraries. Recognising that background is half the skill:
| Noise source | Typical strings |
|---|---|
| PE layout | !This program cannot be run in DOS mode., .text, .rdata, .reloc |
| Import table | DLL and API names (read these in Reading Capabilities from Imports) |
| mingw-w64 runtime | Mingw-w64 runtime failure:, Unknown pseudo relocation protocol version %d., GCC: (GNU) ... |
| MSVC CRT | bad allocation, Unknown exception, locale names, R6016-style runtime errors |
| Statically linked libraries | OpenSSL, zlib or SQLite error messages and version strings |
| Go and Rust | Thousands of runtime strings, package paths, panic messages |
Library strings are not useless: inflate 1.2.13 Copyright tells you zlib is
linked in, and a Go package path such as main.encryptFiles is a gift. But
they are context, not findings.
What you are hunting is author content:
| Category | Examples (lab-safe) | Why it matters |
|---|---|---|
| URLs, domains, IPs | http://update.example.com/lab, 198.51.100.23 | Network IOCs, C2 |
| File paths | C:\ProgramData\LabApp\config.ini, %APPDATA%\Lab\ | Drop locations, host IOCs |
| Registry keys | Software\Microsoft\Windows\CurrentVersion\Run | Persistence |
| Mutex and event names | Global\LabAppMutex | Single-instance check; excellent host IOC |
| User-agents | Mozilla/5.0 (Windows NT 10.0; Win64; x64) | Network detection, often hard-coded and slightly wrong |
| Commands | cmd.exe /c, schtasks /create, vssadmin delete shadows | Behaviour before you run anything |
| Error and log messages | failed to inject, [+] beacon sent | Reveal intent and logic; great YARA material |
| PDB paths | C:\Users\dev\...\Release\Updater.pdb | Clustering and attribution (see From Source to Binary) |
| Format strings | id=%s&os=%d&v=%s | Protocol structure of a C2 check-in |
| API names as text | GetComputerNameA, VirtualAllocEx | Dynamic resolution (see dynamic import resolution) |
| Crypto and encoding artefacts | a 64-character Base64 alphabet, long hex blobs, PEM headers | Something is decoded at runtime |
Grep for the high-value classes first, then read the rest top to bottom:
strings -n 6 sample.bin > ascii.txt
strings -n 6 -e l sample.bin > wide.txt
grep -hiE 'https?://|[0-9]{1,3}(\.[0-9]{1,3}){3}|\.pdb|\\software\\|global\\|mozilla|cmd\.exe|powershell' ascii.txt wide.txtTip: Keep the offsets (
-t x) when a string matters. An offset lets you jump to the string in a hex editor or map it to a section and RVA, and from there to the code that references it in a disassembler.
A good string does double duty: it is an IOC for the SOC and a YARA anchor for detection engineers. Error messages and format strings are often better detection content than URLs, because operators change infrastructure far more often than code. You will turn exactly these into rules in Writing Your First YARA Rules (see also the YARA glossary entry).
Why and how authors hide strings
A binary whose strings list reads like its feature list is easy to detect and easy to analyse. So malware hides the strings that matter, usually with cheap tricks rather than strong crypto:
- Stack strings. The string never exists as a contiguous literal. The code
writes it into a stack buffer one character, or a few characters, at a time
with
movimmediates. The bytes are scattered through instructions, wherestringsdoes not find them. See stack strings. - Single-byte XOR. Every byte is XORed with one key. It is trivial to write, and it removes both the literal text and simple signature matches. Its weakness: zero bytes encode to the key itself, so padding after an encoded string often shows up as a run of one repeated character. See XOR string encryption.
- Multi-byte or rolling XOR. A repeating key or a key that changes per byte. It is still weak, but no longer brute-forceable by trying 255 keys.
- Base64 and custom alphabets. Base64 hides nothing from a human, but it
does defeat naive
grep. A shuffled 64-character alphabet in the strings output is a strong hint of custom Base64. - RC4 and real ciphers. RC4 is short to implement and common in malware configs. Look for a loop that initialises a 256-byte array with 0 to 255 (the key schedule) and a high-entropy blob nearby. AES appears as S-box constants that tools such as capa or signsrch recognise.
- No strings at all. API hashing replaces API names with 32-bit constants, so there is nothing left to decode.
The signs that strings are hidden are indirect: a large binary with almost no
readable author content, printable gibberish near a decoding loop, a blob with
high entropy in .data or a resource, and the
LoadLibraryA and GetProcAddress pair with no API names in sight.
FLOSS: strings the program would build
FLOSS, the FLARE Obfuscated String Solver from Mandiant, extends strings
with emulation. It reports four kinds of strings:
| Type | What it is | How FLOSS finds it |
|---|---|---|
| Static | Plain ASCII and UTF-16LE | Byte scanning, like strings |
| Stack | Built on the stack by mov instructions | Emulates each function and extracts strings from its stack frame |
| Tight | Stack strings decoded in place by a tight loop | Emulates the loop and reads the frame afterwards |
| Decoded | Produced by a decoding function at runtime | Finds candidate decoders and emulates calls to them |
For decoded strings the idea is simple, even though the implementation is not. FLOSS disassembles the file (using vivisect) and ranks functions by features that decoders tend to have: loops, XOR with a non-zero operand, shifts, many callers. For each candidate it emulates the call sites with the arguments the program would pass, then diffs memory before and after. New printable data that appears in the emulated memory is reported as a decoded string. Nothing runs natively, so it is safe to use on real samples inside your lab.
Emulation has limits. Decoders that depend on a key fetched from the network, the environment or a file, heavily obfuscated control flow, or anti-emulation tricks will produce nothing. An empty FLOSS result does not mean "no hidden strings". When FLOSS fails, a debugger breakpoint after the decoding routine is the fallback.
CyberChef: decoding by hand
When you find an encoded blob and its key, CyberChef is the quickest way to try decodings without writing code. Useful operations:
- From Hex, From Base64 (with an alphabet option for custom tables)
- XOR with a known key, and XOR Brute Force with a crib such as
http - RC4 with a key in hex, UTF-8 or Base64
- Magic, which tries likely decodings automatically and ranks the results
Recipes chain, so "From Base64, then RC4 with key X, then Gunzip" is one saved recipe you can reuse on every sample of a family. Run CyberChef offline (it is a single HTML file) when the data comes from a sensitive case.
Lab: four strings, four outcomes
You will build a benign program that stores four strings in four different ways, and see which tools can recover each one.
-
Write the program. The encoded array is
http://update.example.com/labXORed with0x5A:c // strlab.c - four ways to store a string (benign lab program) #include <windows.h> #include <stdio.h> /* 1. Plain ASCII: sits in .rdata as-is */ static const char PLAIN[] = "lab-plain: C:\\ProgramData\\LabApp\\config.ini"; /* 2. UTF-16LE: every character followed by a zero byte */ static const wchar_t WIDE[] = L"lab-wide: Global\\LabAppMutex"; /* 4. Single-byte XOR with key 0x5A */ static unsigned char ENC[] = { 0x32, 0x2e, 0x2e, 0x2a, 0x60, 0x75, 0x75, 0x2f, 0x2a, 0x3e, 0x3b, 0x2e, 0x3f, 0x74, 0x3f, 0x22, 0x3b, 0x37, 0x2a, 0x36, 0x3f, 0x74, 0x39, 0x35, 0x37, 0x75, 0x36, 0x3b, 0x38 }; static void xor_decode(unsigned char *buf, size_t len, unsigned char key) { for (size_t i = 0; i < len; i++) buf[i] ^= key; } int main(void) { /* 3. Stack string: built one character at a time at runtime */ char stack[16]; stack[0] = 'l'; stack[1] = 'a'; stack[2] = 'b'; stack[3] = '-'; stack[4] = 's'; stack[5] = 't'; stack[6] = 'a'; stack[7] = 'c'; stack[8] = 'k'; stack[9] = '-'; stack[10] = 'k'; stack[11] = 'e'; stack[12] = 'y'; stack[13] = '\0'; char url[sizeof(ENC) + 1]; memcpy(url, ENC, sizeof(ENC)); xor_decode((unsigned char *)url, sizeof(ENC), 0x5A); url[sizeof(ENC)] = '\0'; printf("%s\n", PLAIN); wprintf(L"%ls\n", WIDE); printf("%s\n", stack); printf("decoded: %s\n", url); return 0; }bash x86_64-w64-mingw32-gcc -O0 -s -o strlab.exe strlab.cIf you run it in your Windows VM (or under Wine), it prints all four strings in clear text. The program always has the plain text in memory at runtime; the question is what a static look reveals.
-
ASCII strings. Search for the lab markers:
bash strings -n 6 strlab.exe | grep -iE 'lab|http|update'Only
lab-plain: C:\ProgramData\LabApp\config.iniappears. Now read the full output around it. You should finddecoded: %s, a format string that says something is decoded, and a line of printable gibberish that starts with2..*and containsuu/*. That is the XOR-encoded URL, which happens to be printable because lowercase letters XOR0x5Aland in the printable range. -
UTF-16LE strings.
bash strings -n 6 -e l strlab.exelab-wide: Global\LabAppMutexappears, and nothing else from your program. On Windows,strings64.exe -n 6 strlab.exefinds both the plain and wide strings in one pass. -
Look for the stack string.
strings -n 4 strlab.exe | grep -i stackfinds nothing. Disassemblemainand look at how it is built:bash x86_64-w64-mingw32-objdump -d -M intel strlab.exe | grep -m6 'BYTE PTR \[rbp-0x'text c6 45 f0 6c mov BYTE PTR [rbp-0x10],0x6c ; 'l' c6 45 f1 61 mov BYTE PTR [rbp-0xf],0x61 ; 'a' c6 45 f2 62 mov BYTE PTR [rbp-0xe],0x62 ; 'b'Each character is separated from the next by three bytes of opcode and displacement, so no printable run is longer than one character.
-
Run FLOSS. Install it in a virtual environment (
pip install flare-floss) or download a release binary, then ask only for the non-static types:bash floss --only stack tight decoded -- strlab.exeFLOSS reports one stack string,
lab-stack-key, and one decoded string,http://update.example.com/lab, which it recovered by emulatingxor_decode. Run plainfloss strlab.exetoo and confirm that its static section contains both the ASCII and the UTF-16 marker. -
Brute-force single-byte XOR. Write a scanner that tries all 255 keys. It encodes the crib with each key and searches for that, which is much faster than decoding the whole file 255 times:
python # xorscan.py - brute-force single-byte XOR over a file, looking for a crib import re, sys data = open(sys.argv[1], "rb").read() crib = sys.argv[2].encode() if len(sys.argv) > 2 else b"http" printable = re.compile(rb"[\x20-\x7e]{6,}") for key in range(1, 256): target = bytes(b ^ key for b in crib) start = data.find(target) while start != -1: chunk = bytes(b ^ key for b in data[start:start + 128]) m = printable.match(chunk) text = m.group().decode() if m else chunk[:len(crib)].decode() print(f"key=0x{key:02x} offset=0x{start:x} {text}") start = data.find(target, start + 1)bash python3 xorscan.py strlab.exeExpect a single hit with
key=0x5a, in the.datasection, decoding to the URL followed by a fewZcharacters. Those are the zero bytes after the array, XORed with0x5A(Z). This is the zero-byte weakness from earlier. -
Decode it in CyberChef. Copy the 29 bytes at that offset as hex, and apply From Hex then XOR with key
5a. Then try XOR Brute Force with the cribhttpand see the same key come out on top. -
Watch the optimiser. Rebuild with
-O2and repeat steps 2 and 4:bash x86_64-w64-mingw32-gcc -O2 -s -o strlab_O2.exe strlab.c strings -n 4 strlab_O2.exe | grep -iE 'lab|http'With GCC 15 you should see fragments such as
lab-stacand evenhttp://update.exin.rdata. The compiler noticed thatENCis never modified and folded part of the "runtime" decoding into a constant, and it moved the first half of the stack string into.rdataas well. Real malware defeats this withvolatilebuffers or keys computed at runtime. For you, it is a reminder that obfuscation quality varies by build, and that sibling samples of one family can leak what another hides.
Questions to answer: Which of the four strings would a YARA rule written
from strings output alone miss? Why does single-byte XOR leak its key through
zero bytes, and how would a rolling XOR avoid that? What would FLOSS report if
the XOR key were read from a registry value at runtime? Which of your
recovered strings would you hand to the SOC as IOCs, and which would you keep
for detection logic?
Key takeaways
- Extract both ASCII and UTF-16LE strings; Windows programs keep a lot of text in UTF-16. Use a minimum length of about 6 to cut noise.
- Most strings come from the PE layout, imports, the C runtime and linked libraries. Learn that background so the author's strings stand out.
- URLs, paths, registry keys, mutexes, user-agents, commands, error messages, PDB paths and format strings are the high-value classes for IOCs and YARA.
- Stack strings, XOR, Base64, RC4 and API hashing hide text from
strings. Missing author strings in a large binary are a finding in themselves. - FLOSS emulates code to recover stack, tight and decoded strings. CyberChef and a few lines of Python cover simple encodings, and a debugger covers the rest.