Skip to content

Leçon 3.1 · Triage statique· 40 min

Identifying and Hashing Files

Identify a file by its magic bytes rather than its name, then fingerprint it with cryptographic, import, fuzzy and Rich header hashes.

Cette leçon n’est disponible qu’en anglais pour le moment.

Objectifs

  • Identify a file's real type from its magic bytes with file, TrID and Detect It Easy
  • Explain what cryptographic hashes prove, and why one changed byte makes them useless for similarity
  • Use imphash, ssdeep, TLSH and the Rich header hash to cluster related samples
  • Run reputation lookups by hash without leaking confidential files
  • Keep a triage log that other analysts can rely on

Triage starts with two questions that sound trivial: what is this file, and have we seen it before? Get the first one wrong and you open a script in a PE parser, or an executable in a PDF reader. Skip the second and you spend a day reversing something a colleague finished last month.

This lesson opens Module 3. You now know the PE and ELF formats from the inside; from here on you apply that knowledge quickly, in the order a real triage follows. First comes identity: the file's true type, then a set of hashes that each answer a different question.

Names and extensions lie

The file name and extension are chosen by whoever sent the file. Windows uses the extension to decide what to do on a double-click, which is exactly why attackers abuse it:

  • invoice.pdf.exe with "hide extensions for known file types" enabled shows up as invoice.pdf.
  • A right-to-left override character (U+202E) inserted into report[U+202E]fdp.exe makes Explorer display the name as reportexe.pdf.
  • A PE renamed to update.dat or thumbs.db will not run on a click, but a loader can still map it. The extension says nothing about the content.

So the first rule of triage: ignore the name, read the bytes. Record the original name in your notes as evidence, never as a fact about the file.

Magic bytes

Most binary formats start with a fixed signature, the magic number. You met two of them already: MZ for PE files (PE Headers) and \x7fELF (ELF for Malware Analysts). These are the ones you will see most in malware triage:

FormatFirst bytes (hex)ASCIINotes
PE / DOS executable4D 5AMZConfirm with PE\0\0 at e_lfanew
ELF7F 45 4C 46.ELFByte 4 gives 32/64-bit, byte 5 endianness
ZIP (and DOCX, XLSX, JAR, APK)50 4B 03 04PK..Office Open XML is a ZIP; look inside
PDF25 50 44 46 2D%PDF-Readers tolerate junk before it
OLE compound fileD0 CF 11 E0 A1 B1 1A E1Legacy .doc/.xls/.ppt, .msi
7-Zip37 7A BC AF 27 1C7z..'.
RAR52 61 72 21 1A 07Rar!..Followed by 00 (v4) or 01 00 (v5)
Windows shortcut (LNK)4C 00 00 00 01 14 02 00L...Header size, then the shell link CLSID
gzip1F 8BCommon for Linux payloads
Cabinet4D 53 43 46MSCF

A signature is necessary, not sufficient. PK tells you "ZIP container", not whether it is a Word document, a Java archive or an Android app. And many formats, such as scripts, HTA files, batch files and PowerShell, have no magic at all: they are just text.

Warning: Parsers are often more forgiving than signatures suggest. PDF readers accept a header that is not at offset 0, and archive tools will find a ZIP whose central directory sits at the end of another file. A file can be a valid GIF and a valid ZIP at once (a "polyglot"). If a type check and a tool's behaviour disagree, trust neither until you have looked at the bytes.

file and libmagic

On Linux and macOS, file matches the content against the libmagic database of signatures and structural tests:

bash
cp hello.exe invoice.pdf
file invoice.pdf
xxd -l 16 invoice.pdf
text
invoice.pdf: PE32+ executable (console) x86-64 (stripped to external PDB), for MS Windows
00000000: 4d5a 9000 0300 0000 0400 0000 ffff 0000  MZ..............

The extension says PDF, the bytes say PE32+. file -z looks inside compressed files, and file -k keeps going after the first match, which helps with polyglots.

TrID and Detect It Easy

On Windows, two tools do the same job with different strengths:

  • TrID scores a file against thousands of community definitions and prints ranked candidates with percentages. It is useful for obscure formats that libmagic does not know.
  • Detect It Easy (DIE) goes further for executables: besides the format, its signatures name the compiler, linker, packer or protector, such as MSVC, mingw-w64, Go, .NET, UPX or Themida, and it shows per-section entropy. For a PE, DIE is usually the first tool to open. The packer side of it gets its own lesson in Detecting Packers and Entropy.

All three are heuristic. When they disagree, open a hex view and check the signature yourself.

Cryptographic hashes: identity

A cryptographic hash maps any input to a fixed-size digest. Two properties matter to an analyst:

  • Identity. In practice, the same digest means the same bytes, so a hash names a file uniquely and exactly. That is why hashes are the primary key of every sample database, sandbox report and IOC feed.
  • Avalanche. Change one bit anywhere and roughly half the digest bits flip. The new hash tells you nothing about how close the files were.
AlgorithmDigestStatusWhere you still see it
MD5128 bitsCollisions are practical since 2004Older IOC feeds, EDR consoles, imphash
SHA-1160 bitsPractical collision shown in 2017 (SHAttered)Git, legacy tools, VirusTotal
SHA-256256 bitsNo known practical attackThe standard to record and share

Collisions matter when an attacker can craft two files with the same hash, for example to get a benign twin allow-listed. For naming samples and sharing IOCs, MD5 and SHA-1 still work, but record SHA-256 as the canonical identifier and include the others because many tools and feeds still expect them.

powershell
Get-FileHash .\sample.bin -Algorithm SHA256
certutil -hashfile .\sample.bin SHA256
bash
sha256sum sample.bin    # macOS: shasum -a 256 sample.bin

The avalanche property is also the limitation. A file hash is a perfect IOC for this exact file and useless for the next build. Malware operators know it: builders that stamp a unique campaign ID into each copy, or that append random bytes to the overlay (see Resources and Overlays), give every victim a different SHA-256. That is why hashes sit at the bottom of the "pyramid of pain": they cost the attacker nothing to change.

Tip: Hash the file you received and everything you extract from it (the ZIP, the document inside, the dropped EXE), and record the parent-child relationship. Responders searching endpoints need the hash of the file that actually lands on disk, which is rarely the email attachment.

Similarity hashes: relatives

To find related files you need hashes that change a little when the file changes a little. Four are in daily use.

Imphash

You already met imphash in Imports, Exports and the IAT: an MD5 over the normalised import list. It ignores code and data entirely, so it survives changed strings, recompiled code and appended junk, as long as the import list and its order stay the same. It collides for packed files and for tiny or runtime-heavy import tables.

ssdeep: context-triggered piecewise hashing

ssdeep (Kornblum, 2006) splits the file into chunks at boundaries chosen by the content, not at fixed offsets. A rolling hash over a small sliding window triggers a boundary whenever its value hits a condition that depends on a block size. Each chunk is hashed and reduced to a single Base64 character, and the characters are concatenated:

text
192:PFNFU39HUbMetCyMYUdHQNEmeyVxHk/wjE6zCPiMpnHHmyVAgsZguy0HLVDSDRmb:Pu9HXjyVxEpUCNzVBuy0HLuscq
 │   │                                                             │
 │   └ one character per chunk at the block size                   └ same, at double block size
 └ block size

Because boundaries depend on local content, inserting bytes in one place shifts only nearby chunks, and the rest of the signature stays the same. Comparing two signatures gives a score from 0 to 100 based on how many edits turn one string into the other. The limits:

  • Signatures are comparable only when their block sizes are equal or differ by a factor of two. Files of very different sizes always score 0.
  • Small files produce short signatures and unstable scores.
  • Compression, encryption or packing rewrite every byte, so two packed copies of the same program look unrelated.

TLSH

TLSH, the Trend Micro Locality Sensitive Hash, takes a different approach. It slides a 5-byte window across the file, counts byte triplets into 128 buckets, and encodes each bucket's count as a quartile. It adds a header built from the file length and a checksum. Comparing two TLSH digests gives a distance: 0 means near-identical, and it grows without a fixed ceiling as files diverge. Unlike ssdeep, TLSH compares files of different sizes and is harder to fool by reordering content, which is why several large threat-intel platforms index it. It needs at least 50 bytes with enough variety and returns TNULL otherwise.

The Rich header hash

For MSVC-built files, the Rich header records the toolchain that built each object. A hash of its decoded contents groups samples from the same build environment, even across different projects. VirusTotal shows it, and pefile computes one with pe.get_rich_header_hash(). mingw-w64, Go and Delphi produce no Rich header, so the hash is empty, and attackers can copy or forge the header.

Which hash answers which question

HashSurvivesBreaks onBest for
SHA-256NothingAny changeExact identity, IOCs, deduplication
ImphashCode and data editsImport list changes, packingSame codebase, different builds
ssdeepLocal edits, appended dataLarge size change, packingNear-duplicates, same-size variants
TLSHEdits, reordering, size changePacking, encryptionClustering at scale
Rich hashEverything except toolchainNon-MSVC toolchains, forgerySame build machine

No single similarity hash is reliable on its own. Two independent hashes agreeing is a lead; add matching strings, code or infrastructure and it becomes a finding.

Reputation lookups, done safely

With hashes in hand, ask whether the world already knows this file: your own EDR and sample store first, then public services such as VirusTotal and MalwareBazaar. Search by hash; do not upload.

The difference matters. A hash search reveals almost nothing: the digest cannot be reversed into the file. An upload hands over the whole file, and on most public services it becomes available to other subscribers who can download submitted files. That creates two separate risks:

  • Data leak. A phishing document tailored to your organisation may contain real names, internal project names or a stolen spreadsheet. Malware that carries a victim-specific config may contain your domain, proxy settings or even credentials. Uploading it publishes that to strangers.
  • Tipping off the attacker. Operators of targeted campaigns watch public services for their own hashes. A first-seen upload of a one-off implant tells them they have been discovered, often before your incident response is ready.

Warning: The same caution applies to URLs and domains you extract. Asking a public scanner to "scan" a URL makes it fetch that URL from its own infrastructure, which the operator may see. Look up the defanged indicator in historical data instead, and follow your organisation's sharing policy (for example the Traffic Light Protocol) before anything leaves the building.

If a hash lookup returns nothing, that is also a result: the sample is new, targeted or freshly rebuilt. Note it and move on to the rest of the triage.

Keep a triage log

Every file you touch gets a row, recorded before you start experimenting with it. A spreadsheet is enough; the discipline is what matters:

text
date        sha256 (full)   md5 / sha1   size    type (magic + DIE)        original name / source
imphash     ssdeep          tlsh         rich    parent sha256             lookup result + date
analyst     notes (defanged indicators, next steps)

Record the lookup date: "0 detections" means something different on day one and a month later. Record the parent hash so an extracted payload can always be traced to the email or archive it came from. And keep the log with the case, not in your head, so the next analyst can pick up where you left off.

Lab: which hashes see the family?

You will build one benign program, derive small variants from it, and measure how each hash reacts. The commands use mingw-w64 on Linux or macOS; any Python 3.9+ works for the script.

  1. Write the base program:

    c
    // triage.c - benign lab program
    #include <windows.h>
    #include <stdio.h>
    
    static const char *BANNER = "lab build A: checking for updates";
    static const char *SERVER = "http://update.example.com/lab";
    
    int main(void) {
        SYSTEMTIME st;
        GetLocalTime(&st);
        printf("%s\n", BANNER);
        printf("would contact %s at %02d:%02d\n", SERVER, st.wHour, st.wMinute);
        Sleep(100);
        return 0;
    }
  2. Build the variants. Each differs from v1_base.exe in one controlled way:

    bash
    x86_64-w64-mingw32-gcc -O0 -s -o v1_base.exe triage.c
    x86_64-w64-mingw32-gcc -O0 -s -o v1_again.exe triage.c           # same source, rebuilt
    sed 's/lab build A/lab build B/' triage.c > triage_b.c
    x86_64-w64-mingw32-gcc -O0 -s -o v2_string.exe triage_b.c         # one character changed
    x86_64-w64-mingw32-gcc -O2 -s -o v3_O2.exe triage.c               # different optimisation
    x86_64-w64-mingw32-gcc -O0    -o v4_symbols.exe triage.c          # not stripped
    sed 's/Sleep(100);/Sleep(100); printf("%lu\\n", GetTickCount());/' triage.c > triage_c.c
    x86_64-w64-mingw32-gcc -O0 -s -o v6_import.exe triage_c.c         # one extra import
    cp v1_base.exe v5_flipped.exe
    printf '\x01' | dd of=v5_flipped.exe bs=1 \
        seek=$(( $(wc -c < v1_base.exe) - 16 )) conv=notrunc          # one byte near the end
  3. Install the libraries in a virtual environment. ppdeep is a pure-Python ssdeep implementation; python-ssdeep works too if libfuzzy is installed.

    bash
    python3 -m venv venv && . venv/bin/activate
    pip install pefile ppdeep py-tlsh
  4. Compute and compare every hash against the base build:

    python
    # hashes.py - compare identity and similarity hashes across variants
    import sys, hashlib, pefile, ppdeep, tlsh
    
    def row(path):
        data = open(path, "rb").read()
        pe = pefile.PE(data=data)
        return {"file": path,
                "sha256": hashlib.sha256(data).hexdigest(),
                "imphash": pe.get_imphash(),
                "ssdeep": ppdeep.hash(data),
                "tlsh": tlsh.hash(data)}
    
    rows = [row(p) for p in sys.argv[1:]]
    base = rows[0]
    print(f"{'vs ' + base['file']:<18}{'sha256':>8}{'imphash':>9}{'ssdeep':>8}{'tlsh':>6}")
    for r in rows[1:]:
        same = lambda k: "same" if r[k] == base[k] else "diff"
        print(f"{r['file']:<18}{same('sha256'):>8}{same('imphash'):>9}"
              f"{ppdeep.compare(base['ssdeep'], r['ssdeep']):>8}"
              f"{tlsh.diff(base['tlsh'], r['tlsh']):>6}")
    bash
    python3 hashes.py v1_base.exe v1_again.exe v2_string.exe v5_flipped.exe \
        v3_O2.exe v6_import.exe v4_symbols.exe

    A run with mingw-w64 GCC 15 produced the table below. Your numbers will differ with the compiler version, and the TLSH distances move by a few points between otherwise identical builds, but the pattern should hold. ssdeep is a similarity score (higher is closer); TLSH is a distance (lower is closer).

    text
    vs v1_base.exe      sha256  imphash  ssdeep  tlsh
    v1_again.exe          diff     same     100     1
    v2_string.exe         diff     same      99     3
    v5_flipped.exe        diff     same     100     1
    v3_O2.exe             diff     same      65    32
    v6_import.exe         diff     diff      71    48
    v4_symbols.exe        diff     same       0   511
  5. Explain v1_again.exe. Same source, same compiler, yet a different SHA-256. Find the bytes that differ:

    bash
    cmp -l v1_base.exe v1_again.exe

    The differences sit in the COFF TimeDateStamp and the optional header CheckSum (see PE Headers). Every build is a new file hash, even with no code change at all.

  6. Check the Rich header hash. Run python3 -c "import pefile; print(repr(pefile.PE('v1_base.exe').get_rich_header_hash()))". A mingw build returns an empty string. If you have MSVC, build triage.c with cl twice from different Visual Studio versions and compare.

  7. Identify every variant with file, and if you have a Windows VM, drop them into Detect It Easy and TrID. Rename one to .pdf and confirm that none of the tools care.

  8. Start your triage log with one row per variant.

Questions to answer: Why does v4_symbols.exe score 0 in ssdeep but still get a finite TLSH distance? Which hash would still link v6_import.exe to the base if the author also packed it with UPX? Which two hashes would you publish as IOCs for this "family", and which would you keep internal as clustering pivots? What would happen to all five hashes if the builder appended 4 KB of random bytes to the overlay?

Key takeaways

  • Identify files by magic bytes and structure, never by name or extension; file, TrID and DIE are heuristics you confirm with a hex view.
  • SHA-256 names one exact file. A rebuild, a timestamp or a single flipped byte gives a new hash, so file hashes are precise but brittle IOCs.
  • Imphash, ssdeep, TLSH and the Rich header hash each survive different changes. Use two or more together to cluster, and packing defeats most of them.
  • Look up hashes, do not upload files: uploads can leak victim data and warn the attacker.
  • Log every file with its hashes, type, source, parent and lookup date before you go deeper.