Skip to content

Leçon 3.2 · Triage statique· 45 min

Strings and Obfuscated Strings

Extract ASCII and UTF-16 strings, separate indicators from runtime noise, and recover stack, XOR and decoded strings with FLOSS and CyberChef.

Cette leçon n’est disponible qu’en anglais pour le moment.

Objectifs

  • Extract ASCII and UTF-16LE strings with GNU strings, Sysinternals strings and FLOSS
  • Separate analyst-relevant strings from compiler and runtime noise
  • Recognise stack strings, single-byte XOR, Base64 and RC4 as ways of hiding strings
  • Explain at a high level how FLOSS recovers stack, tight and decoded strings
  • Brute-force single-byte XOR with a known-plaintext crib

After identity and hashes, the next triage step is reading the text the file carries. A C2 URL, a mutex name, a ransom note or a developer's PDB path can tell you more in thirty seconds than an hour in a disassembler. That is also why malware authors hide their strings, and why noticing missing strings is as important as reading the ones that are there.

You already used strings in earlier labs. This lesson makes it a method: which encodings to search, how to cut the noise, what to look for, and what to do when the interesting text is not in the file in readable form.

What a "string" is to a tool

A strings extractor knows nothing about the file format. It scans the raw bytes for runs of printable characters at least N long and prints each run. Two encodings matter on Windows:

text
"Lab" as ASCII / UTF-8        4C 61 62 00
"Lab" as UTF-16LE (wchar_t)   4C 00 61 00 62 00 00 00

Windows APIs ending in W (CreateFileW, RegSetValueExW) take UTF-16LE, and so do L"..." literals in C, .NET string tables, resources and the version-info block. An ASCII scanner sees L, a zero byte, a, a zero byte and gives up, because each run is only one character long. Always search both encodings.

ToolASCIIUTF-16LENotes
GNU strings (binutils)default-e lDefault minimum length 4; -t x prints hex offsets
Sysinternals strings-a only-u onlySearches both by default; minimum length 3; -o prints offsets
FLOSSyesyesAlso recovers stack, tight and decoded strings

On macOS the built-in strings is Apple's and has no -e option. Use the one from binutils or mingw-w64 (x86_64-w64-mingw32-strings).

Minimum length and noise

The minimum length is a noise dial. Machine code and compressed data contain short printable runs by pure chance. On a small mingw-w64 test program, the ASCII run count drops from about 200 at -n 4 to about 115 at -n 6 and under 100 at -n 8, and none of the lost strings were interesting. Start at 6, and drop to 4 only when you are hunting something specific, such as short mutex names or two-letter commands.

Random-looking short runs such as AWAVAUATUWVSH or [^_]A\A]A^A_ are x86-64 instructions: the REX-prefixed push and pop of callee-saved registers in function prologues and epilogues happen to be printable. Learn to recognise them and skip them.

Triage: signal versus noise

Most strings in a binary were not written by the author. They come from the compiler, the C runtime and statically linked libraries. Recognising that background is half the skill:

Noise sourceTypical strings
PE layout!This program cannot be run in DOS mode., .text, .rdata, .reloc
Import tableDLL and API names (read these in Reading Capabilities from Imports)
mingw-w64 runtimeMingw-w64 runtime failure:, Unknown pseudo relocation protocol version %d., GCC: (GNU) ...
MSVC CRTbad allocation, Unknown exception, locale names, R6016-style runtime errors
Statically linked librariesOpenSSL, zlib or SQLite error messages and version strings
Go and RustThousands of runtime strings, package paths, panic messages

Library strings are not useless: inflate 1.2.13 Copyright tells you zlib is linked in, and a Go package path such as main.encryptFiles is a gift. But they are context, not findings.

What you are hunting is author content:

CategoryExamples (lab-safe)Why it matters
URLs, domains, IPshttp://update.example.com/lab, 198.51.100.23Network IOCs, C2
File pathsC:\ProgramData\LabApp\config.ini, %APPDATA%\Lab\Drop locations, host IOCs
Registry keysSoftware\Microsoft\Windows\CurrentVersion\RunPersistence
Mutex and event namesGlobal\LabAppMutexSingle-instance check; excellent host IOC
User-agentsMozilla/5.0 (Windows NT 10.0; Win64; x64)Network detection, often hard-coded and slightly wrong
Commandscmd.exe /c, schtasks /create, vssadmin delete shadowsBehaviour before you run anything
Error and log messagesfailed to inject, [+] beacon sentReveal intent and logic; great YARA material
PDB pathsC:\Users\dev\...\Release\Updater.pdbClustering and attribution (see From Source to Binary)
Format stringsid=%s&os=%d&v=%sProtocol structure of a C2 check-in
API names as textGetComputerNameA, VirtualAllocExDynamic resolution (see dynamic import resolution)
Crypto and encoding artefactsa 64-character Base64 alphabet, long hex blobs, PEM headersSomething is decoded at runtime

Grep for the high-value classes first, then read the rest top to bottom:

bash
strings -n 6 sample.bin > ascii.txt
strings -n 6 -e l sample.bin > wide.txt
grep -hiE 'https?://|[0-9]{1,3}(\.[0-9]{1,3}){3}|\.pdb|\\software\\|global\\|mozilla|cmd\.exe|powershell' ascii.txt wide.txt

Tip: Keep the offsets (-t x) when a string matters. An offset lets you jump to the string in a hex editor or map it to a section and RVA, and from there to the code that references it in a disassembler.

A good string does double duty: it is an IOC for the SOC and a YARA anchor for detection engineers. Error messages and format strings are often better detection content than URLs, because operators change infrastructure far more often than code. You will turn exactly these into rules in Writing Your First YARA Rules (see also the YARA glossary entry).

Why and how authors hide strings

A binary whose strings list reads like its feature list is easy to detect and easy to analyse. So malware hides the strings that matter, usually with cheap tricks rather than strong crypto:

  • Stack strings. The string never exists as a contiguous literal. The code writes it into a stack buffer one character, or a few characters, at a time with mov immediates. The bytes are scattered through instructions, where strings does not find them. See stack strings.
  • Single-byte XOR. Every byte is XORed with one key. It is trivial to write, and it removes both the literal text and simple signature matches. Its weakness: zero bytes encode to the key itself, so padding after an encoded string often shows up as a run of one repeated character. See XOR string encryption.
  • Multi-byte or rolling XOR. A repeating key or a key that changes per byte. It is still weak, but no longer brute-forceable by trying 255 keys.
  • Base64 and custom alphabets. Base64 hides nothing from a human, but it does defeat naive grep. A shuffled 64-character alphabet in the strings output is a strong hint of custom Base64.
  • RC4 and real ciphers. RC4 is short to implement and common in malware configs. Look for a loop that initialises a 256-byte array with 0 to 255 (the key schedule) and a high-entropy blob nearby. AES appears as S-box constants that tools such as capa or signsrch recognise.
  • No strings at all. API hashing replaces API names with 32-bit constants, so there is nothing left to decode.

The signs that strings are hidden are indirect: a large binary with almost no readable author content, printable gibberish near a decoding loop, a blob with high entropy in .data or a resource, and the LoadLibraryA and GetProcAddress pair with no API names in sight.

FLOSS: strings the program would build

FLOSS, the FLARE Obfuscated String Solver from Mandiant, extends strings with emulation. It reports four kinds of strings:

TypeWhat it isHow FLOSS finds it
StaticPlain ASCII and UTF-16LEByte scanning, like strings
StackBuilt on the stack by mov instructionsEmulates each function and extracts strings from its stack frame
TightStack strings decoded in place by a tight loopEmulates the loop and reads the frame afterwards
DecodedProduced by a decoding function at runtimeFinds candidate decoders and emulates calls to them

For decoded strings the idea is simple, even though the implementation is not. FLOSS disassembles the file (using vivisect) and ranks functions by features that decoders tend to have: loops, XOR with a non-zero operand, shifts, many callers. For each candidate it emulates the call sites with the arguments the program would pass, then diffs memory before and after. New printable data that appears in the emulated memory is reported as a decoded string. Nothing runs natively, so it is safe to use on real samples inside your lab.

Emulation has limits. Decoders that depend on a key fetched from the network, the environment or a file, heavily obfuscated control flow, or anti-emulation tricks will produce nothing. An empty FLOSS result does not mean "no hidden strings". When FLOSS fails, a debugger breakpoint after the decoding routine is the fallback.

CyberChef: decoding by hand

When you find an encoded blob and its key, CyberChef is the quickest way to try decodings without writing code. Useful operations:

  • From Hex, From Base64 (with an alphabet option for custom tables)
  • XOR with a known key, and XOR Brute Force with a crib such as http
  • RC4 with a key in hex, UTF-8 or Base64
  • Magic, which tries likely decodings automatically and ranks the results

Recipes chain, so "From Base64, then RC4 with key X, then Gunzip" is one saved recipe you can reuse on every sample of a family. Run CyberChef offline (it is a single HTML file) when the data comes from a sensitive case.

Lab: four strings, four outcomes

You will build a benign program that stores four strings in four different ways, and see which tools can recover each one.

  1. Write the program. The encoded array is http://update.example.com/lab XORed with 0x5A:

    c
    // strlab.c - four ways to store a string (benign lab program)
    #include <windows.h>
    #include <stdio.h>
    
    /* 1. Plain ASCII: sits in .rdata as-is */
    static const char PLAIN[] = "lab-plain: C:\\ProgramData\\LabApp\\config.ini";
    
    /* 2. UTF-16LE: every character followed by a zero byte */
    static const wchar_t WIDE[] = L"lab-wide: Global\\LabAppMutex";
    
    /* 4. Single-byte XOR with key 0x5A */
    static unsigned char ENC[] = {
        0x32, 0x2e, 0x2e, 0x2a, 0x60, 0x75, 0x75, 0x2f, 0x2a, 0x3e, 0x3b, 0x2e,
        0x3f, 0x74, 0x3f, 0x22, 0x3b, 0x37, 0x2a, 0x36, 0x3f, 0x74, 0x39, 0x35,
        0x37, 0x75, 0x36, 0x3b, 0x38
    };
    
    static void xor_decode(unsigned char *buf, size_t len, unsigned char key) {
        for (size_t i = 0; i < len; i++)
            buf[i] ^= key;
    }
    
    int main(void) {
        /* 3. Stack string: built one character at a time at runtime */
        char stack[16];
        stack[0] = 'l'; stack[1] = 'a'; stack[2] = 'b'; stack[3] = '-';
        stack[4] = 's'; stack[5] = 't'; stack[6] = 'a'; stack[7] = 'c';
        stack[8] = 'k'; stack[9] = '-'; stack[10] = 'k'; stack[11] = 'e';
        stack[12] = 'y'; stack[13] = '\0';
    
        char url[sizeof(ENC) + 1];
        memcpy(url, ENC, sizeof(ENC));
        xor_decode((unsigned char *)url, sizeof(ENC), 0x5A);
        url[sizeof(ENC)] = '\0';
    
        printf("%s\n", PLAIN);
        wprintf(L"%ls\n", WIDE);
        printf("%s\n", stack);
        printf("decoded: %s\n", url);
        return 0;
    }
    bash
    x86_64-w64-mingw32-gcc -O0 -s -o strlab.exe strlab.c

    If you run it in your Windows VM (or under Wine), it prints all four strings in clear text. The program always has the plain text in memory at runtime; the question is what a static look reveals.

  2. ASCII strings. Search for the lab markers:

    bash
    strings -n 6 strlab.exe | grep -iE 'lab|http|update'

    Only lab-plain: C:\ProgramData\LabApp\config.ini appears. Now read the full output around it. You should find decoded: %s, a format string that says something is decoded, and a line of printable gibberish that starts with 2..* and contains uu/*. That is the XOR-encoded URL, which happens to be printable because lowercase letters XOR 0x5A land in the printable range.

  3. UTF-16LE strings.

    bash
    strings -n 6 -e l strlab.exe

    lab-wide: Global\LabAppMutex appears, and nothing else from your program. On Windows, strings64.exe -n 6 strlab.exe finds both the plain and wide strings in one pass.

  4. Look for the stack string. strings -n 4 strlab.exe | grep -i stack finds nothing. Disassemble main and look at how it is built:

    bash
    x86_64-w64-mingw32-objdump -d -M intel strlab.exe | grep -m6 'BYTE PTR \[rbp-0x'
    text
    c6 45 f0 6c    mov    BYTE PTR [rbp-0x10],0x6c    ; 'l'
    c6 45 f1 61    mov    BYTE PTR [rbp-0xf],0x61     ; 'a'
    c6 45 f2 62    mov    BYTE PTR [rbp-0xe],0x62     ; 'b'

    Each character is separated from the next by three bytes of opcode and displacement, so no printable run is longer than one character.

  5. Run FLOSS. Install it in a virtual environment (pip install flare-floss) or download a release binary, then ask only for the non-static types:

    bash
    floss --only stack tight decoded -- strlab.exe

    FLOSS reports one stack string, lab-stack-key, and one decoded string, http://update.example.com/lab, which it recovered by emulating xor_decode. Run plain floss strlab.exe too and confirm that its static section contains both the ASCII and the UTF-16 marker.

  6. Brute-force single-byte XOR. Write a scanner that tries all 255 keys. It encodes the crib with each key and searches for that, which is much faster than decoding the whole file 255 times:

    python
    # xorscan.py - brute-force single-byte XOR over a file, looking for a crib
    import re, sys
    
    data = open(sys.argv[1], "rb").read()
    crib = sys.argv[2].encode() if len(sys.argv) > 2 else b"http"
    printable = re.compile(rb"[\x20-\x7e]{6,}")
    
    for key in range(1, 256):
        target = bytes(b ^ key for b in crib)
        start = data.find(target)
        while start != -1:
            chunk = bytes(b ^ key for b in data[start:start + 128])
            m = printable.match(chunk)
            text = m.group().decode() if m else chunk[:len(crib)].decode()
            print(f"key=0x{key:02x}  offset=0x{start:x}  {text}")
            start = data.find(target, start + 1)
    bash
    python3 xorscan.py strlab.exe

    Expect a single hit with key=0x5a, in the .data section, decoding to the URL followed by a few Z characters. Those are the zero bytes after the array, XORed with 0x5A (Z). This is the zero-byte weakness from earlier.

  7. Decode it in CyberChef. Copy the 29 bytes at that offset as hex, and apply From Hex then XOR with key 5a. Then try XOR Brute Force with the crib http and see the same key come out on top.

  8. Watch the optimiser. Rebuild with -O2 and repeat steps 2 and 4:

    bash
    x86_64-w64-mingw32-gcc -O2 -s -o strlab_O2.exe strlab.c
    strings -n 4 strlab_O2.exe | grep -iE 'lab|http'

    With GCC 15 you should see fragments such as lab-stac and even http://update.ex in .rdata. The compiler noticed that ENC is never modified and folded part of the "runtime" decoding into a constant, and it moved the first half of the stack string into .rdata as well. Real malware defeats this with volatile buffers or keys computed at runtime. For you, it is a reminder that obfuscation quality varies by build, and that sibling samples of one family can leak what another hides.

Questions to answer: Which of the four strings would a YARA rule written from strings output alone miss? Why does single-byte XOR leak its key through zero bytes, and how would a rolling XOR avoid that? What would FLOSS report if the XOR key were read from a registry value at runtime? Which of your recovered strings would you hand to the SOC as IOCs, and which would you keep for detection logic?

Key takeaways

  • Extract both ASCII and UTF-16LE strings; Windows programs keep a lot of text in UTF-16. Use a minimum length of about 6 to cut noise.
  • Most strings come from the PE layout, imports, the C runtime and linked libraries. Learn that background so the author's strings stand out.
  • URLs, paths, registry keys, mutexes, user-agents, commands, error messages, PDB paths and format strings are the high-value classes for IOCs and YARA.
  • Stack strings, XOR, Base64, RC4 and API hashing hide text from strings. Missing author strings in a large binary are a finding in themselves.
  • FLOSS emulates code to recover stack, tight and decoded strings. CyberChef and a few lines of Python cover simple encodings, and a debugger covers the rest.