Leçon 3.1 · Triage statique· 40 min
Identifying and Hashing Files
Identify a file by its magic bytes rather than its name, then fingerprint it with cryptographic, import, fuzzy and Rich header hashes.
Cette leçon n’est disponible qu’en anglais pour le moment.
Objectifs
- Identify a file's real type from its magic bytes with file, TrID and Detect It Easy
- Explain what cryptographic hashes prove, and why one changed byte makes them useless for similarity
- Use imphash, ssdeep, TLSH and the Rich header hash to cluster related samples
- Run reputation lookups by hash without leaking confidential files
- Keep a triage log that other analysts can rely on
Triage starts with two questions that sound trivial: what is this file, and have we seen it before? Get the first one wrong and you open a script in a PE parser, or an executable in a PDF reader. Skip the second and you spend a day reversing something a colleague finished last month.
This lesson opens Module 3. You now know the PE and ELF formats from the inside; from here on you apply that knowledge quickly, in the order a real triage follows. First comes identity: the file's true type, then a set of hashes that each answer a different question.
Names and extensions lie
The file name and extension are chosen by whoever sent the file. Windows uses the extension to decide what to do on a double-click, which is exactly why attackers abuse it:
invoice.pdf.exewith "hide extensions for known file types" enabled shows up asinvoice.pdf.- A right-to-left override character (U+202E) inserted into
report[U+202E]fdp.exemakes Explorer display the name asreportexe.pdf. - A PE renamed to
update.datorthumbs.dbwill not run on a click, but a loader can still map it. The extension says nothing about the content.
So the first rule of triage: ignore the name, read the bytes. Record the original name in your notes as evidence, never as a fact about the file.
Magic bytes
Most binary formats start with a fixed signature, the magic number. You met
two of them already: MZ for PE files (PE Headers) and
\x7fELF (ELF for Malware Analysts).
These are the ones you will see most in malware triage:
| Format | First bytes (hex) | ASCII | Notes |
|---|---|---|---|
| PE / DOS executable | 4D 5A | MZ | Confirm with PE\0\0 at e_lfanew |
| ELF | 7F 45 4C 46 | .ELF | Byte 4 gives 32/64-bit, byte 5 endianness |
| ZIP (and DOCX, XLSX, JAR, APK) | 50 4B 03 04 | PK.. | Office Open XML is a ZIP; look inside |
25 50 44 46 2D | %PDF- | Readers tolerate junk before it | |
| OLE compound file | D0 CF 11 E0 A1 B1 1A E1 | Legacy .doc/.xls/.ppt, .msi | |
| 7-Zip | 37 7A BC AF 27 1C | 7z..'. | |
| RAR | 52 61 72 21 1A 07 | Rar!.. | Followed by 00 (v4) or 01 00 (v5) |
| Windows shortcut (LNK) | 4C 00 00 00 01 14 02 00 | L... | Header size, then the shell link CLSID |
| gzip | 1F 8B | Common for Linux payloads | |
| Cabinet | 4D 53 43 46 | MSCF |
A signature is necessary, not sufficient. PK tells you "ZIP container", not
whether it is a Word document, a Java archive or an Android app. And many
formats, such as scripts, HTA files, batch files and PowerShell, have no magic at
all: they are just text.
Warning: Parsers are often more forgiving than signatures suggest. PDF readers accept a header that is not at offset 0, and archive tools will find a ZIP whose central directory sits at the end of another file. A file can be a valid GIF and a valid ZIP at once (a "polyglot"). If a type check and a tool's behaviour disagree, trust neither until you have looked at the bytes.
file and libmagic
On Linux and macOS, file matches the content against the libmagic database of
signatures and structural tests:
cp hello.exe invoice.pdf
file invoice.pdf
xxd -l 16 invoice.pdfinvoice.pdf: PE32+ executable (console) x86-64 (stripped to external PDB), for MS Windows
00000000: 4d5a 9000 0300 0000 0400 0000 ffff 0000 MZ..............The extension says PDF, the bytes say PE32+. file -z looks inside compressed
files, and file -k keeps going after the first match, which helps with
polyglots.
TrID and Detect It Easy
On Windows, two tools do the same job with different strengths:
- TrID scores a file against thousands of community definitions and prints ranked candidates with percentages. It is useful for obscure formats that libmagic does not know.
- Detect It Easy (DIE) goes further for executables: besides the format, its signatures name the compiler, linker, packer or protector, such as MSVC, mingw-w64, Go, .NET, UPX or Themida, and it shows per-section entropy. For a PE, DIE is usually the first tool to open. The packer side of it gets its own lesson in Detecting Packers and Entropy.
All three are heuristic. When they disagree, open a hex view and check the signature yourself.
Cryptographic hashes: identity
A cryptographic hash maps any input to a fixed-size digest. Two properties matter to an analyst:
- Identity. In practice, the same digest means the same bytes, so a hash names a file uniquely and exactly. That is why hashes are the primary key of every sample database, sandbox report and IOC feed.
- Avalanche. Change one bit anywhere and roughly half the digest bits flip. The new hash tells you nothing about how close the files were.
| Algorithm | Digest | Status | Where you still see it |
|---|---|---|---|
| MD5 | 128 bits | Collisions are practical since 2004 | Older IOC feeds, EDR consoles, imphash |
| SHA-1 | 160 bits | Practical collision shown in 2017 (SHAttered) | Git, legacy tools, VirusTotal |
| SHA-256 | 256 bits | No known practical attack | The standard to record and share |
Collisions matter when an attacker can craft two files with the same hash, for example to get a benign twin allow-listed. For naming samples and sharing IOCs, MD5 and SHA-1 still work, but record SHA-256 as the canonical identifier and include the others because many tools and feeds still expect them.
Get-FileHash .\sample.bin -Algorithm SHA256
certutil -hashfile .\sample.bin SHA256sha256sum sample.bin # macOS: shasum -a 256 sample.binThe avalanche property is also the limitation. A file hash is a perfect IOC for this exact file and useless for the next build. Malware operators know it: builders that stamp a unique campaign ID into each copy, or that append random bytes to the overlay (see Resources and Overlays), give every victim a different SHA-256. That is why hashes sit at the bottom of the "pyramid of pain": they cost the attacker nothing to change.
Tip: Hash the file you received and everything you extract from it (the ZIP, the document inside, the dropped EXE), and record the parent-child relationship. Responders searching endpoints need the hash of the file that actually lands on disk, which is rarely the email attachment.
Similarity hashes: relatives
To find related files you need hashes that change a little when the file changes a little. Four are in daily use.
Imphash
You already met imphash in Imports, Exports and the IAT: an MD5 over the normalised import list. It ignores code and data entirely, so it survives changed strings, recompiled code and appended junk, as long as the import list and its order stay the same. It collides for packed files and for tiny or runtime-heavy import tables.
ssdeep: context-triggered piecewise hashing
ssdeep (Kornblum, 2006) splits the file into chunks at boundaries chosen by the content, not at fixed offsets. A rolling hash over a small sliding window triggers a boundary whenever its value hits a condition that depends on a block size. Each chunk is hashed and reduced to a single Base64 character, and the characters are concatenated:
192:PFNFU39HUbMetCyMYUdHQNEmeyVxHk/wjE6zCPiMpnHHmyVAgsZguy0HLVDSDRmb:Pu9HXjyVxEpUCNzVBuy0HLuscq
│ │ │
│ └ one character per chunk at the block size └ same, at double block size
└ block sizeBecause boundaries depend on local content, inserting bytes in one place shifts only nearby chunks, and the rest of the signature stays the same. Comparing two signatures gives a score from 0 to 100 based on how many edits turn one string into the other. The limits:
- Signatures are comparable only when their block sizes are equal or differ by a factor of two. Files of very different sizes always score 0.
- Small files produce short signatures and unstable scores.
- Compression, encryption or packing rewrite every byte, so two packed copies of the same program look unrelated.
TLSH
TLSH, the Trend Micro Locality Sensitive Hash, takes a different approach.
It slides a 5-byte window across the file, counts byte triplets into 128
buckets, and encodes each bucket's count as a quartile. It adds a header built
from the file length and a checksum. Comparing two TLSH digests gives a
distance: 0 means near-identical, and it grows without a fixed ceiling as
files diverge. Unlike ssdeep, TLSH compares files of different sizes and is
harder to fool by reordering content, which is why several large threat-intel
platforms index it. It needs at least 50 bytes with enough variety and returns
TNULL otherwise.
The Rich header hash
For MSVC-built files, the Rich header records the
toolchain that built each object. A hash of its decoded contents groups samples
from the same build environment, even across different projects. VirusTotal
shows it, and pefile computes one with pe.get_rich_header_hash(). mingw-w64,
Go and Delphi produce no Rich header, so the hash is empty, and attackers can
copy or forge the header.
Which hash answers which question
| Hash | Survives | Breaks on | Best for |
|---|---|---|---|
| SHA-256 | Nothing | Any change | Exact identity, IOCs, deduplication |
| Imphash | Code and data edits | Import list changes, packing | Same codebase, different builds |
| ssdeep | Local edits, appended data | Large size change, packing | Near-duplicates, same-size variants |
| TLSH | Edits, reordering, size change | Packing, encryption | Clustering at scale |
| Rich hash | Everything except toolchain | Non-MSVC toolchains, forgery | Same build machine |
No single similarity hash is reliable on its own. Two independent hashes agreeing is a lead; add matching strings, code or infrastructure and it becomes a finding.
Reputation lookups, done safely
With hashes in hand, ask whether the world already knows this file: your own EDR and sample store first, then public services such as VirusTotal and MalwareBazaar. Search by hash; do not upload.
The difference matters. A hash search reveals almost nothing: the digest cannot be reversed into the file. An upload hands over the whole file, and on most public services it becomes available to other subscribers who can download submitted files. That creates two separate risks:
- Data leak. A phishing document tailored to your organisation may contain real names, internal project names or a stolen spreadsheet. Malware that carries a victim-specific config may contain your domain, proxy settings or even credentials. Uploading it publishes that to strangers.
- Tipping off the attacker. Operators of targeted campaigns watch public services for their own hashes. A first-seen upload of a one-off implant tells them they have been discovered, often before your incident response is ready.
Warning: The same caution applies to URLs and domains you extract. Asking a public scanner to "scan" a URL makes it fetch that URL from its own infrastructure, which the operator may see. Look up the defanged indicator in historical data instead, and follow your organisation's sharing policy (for example the Traffic Light Protocol) before anything leaves the building.
If a hash lookup returns nothing, that is also a result: the sample is new, targeted or freshly rebuilt. Note it and move on to the rest of the triage.
Keep a triage log
Every file you touch gets a row, recorded before you start experimenting with it. A spreadsheet is enough; the discipline is what matters:
date sha256 (full) md5 / sha1 size type (magic + DIE) original name / source
imphash ssdeep tlsh rich parent sha256 lookup result + date
analyst notes (defanged indicators, next steps)Record the lookup date: "0 detections" means something different on day one and a month later. Record the parent hash so an extracted payload can always be traced to the email or archive it came from. And keep the log with the case, not in your head, so the next analyst can pick up where you left off.
Lab: which hashes see the family?
You will build one benign program, derive small variants from it, and measure how each hash reacts. The commands use mingw-w64 on Linux or macOS; any Python 3.9+ works for the script.
-
Write the base program:
c // triage.c - benign lab program #include <windows.h> #include <stdio.h> static const char *BANNER = "lab build A: checking for updates"; static const char *SERVER = "http://update.example.com/lab"; int main(void) { SYSTEMTIME st; GetLocalTime(&st); printf("%s\n", BANNER); printf("would contact %s at %02d:%02d\n", SERVER, st.wHour, st.wMinute); Sleep(100); return 0; } -
Build the variants. Each differs from
v1_base.exein one controlled way:bash x86_64-w64-mingw32-gcc -O0 -s -o v1_base.exe triage.c x86_64-w64-mingw32-gcc -O0 -s -o v1_again.exe triage.c # same source, rebuilt sed 's/lab build A/lab build B/' triage.c > triage_b.c x86_64-w64-mingw32-gcc -O0 -s -o v2_string.exe triage_b.c # one character changed x86_64-w64-mingw32-gcc -O2 -s -o v3_O2.exe triage.c # different optimisation x86_64-w64-mingw32-gcc -O0 -o v4_symbols.exe triage.c # not stripped sed 's/Sleep(100);/Sleep(100); printf("%lu\\n", GetTickCount());/' triage.c > triage_c.c x86_64-w64-mingw32-gcc -O0 -s -o v6_import.exe triage_c.c # one extra import cp v1_base.exe v5_flipped.exe printf '\x01' | dd of=v5_flipped.exe bs=1 \ seek=$(( $(wc -c < v1_base.exe) - 16 )) conv=notrunc # one byte near the end -
Install the libraries in a virtual environment.
ppdeepis a pure-Python ssdeep implementation;python-ssdeepworks too if libfuzzy is installed.bash python3 -m venv venv && . venv/bin/activate pip install pefile ppdeep py-tlsh -
Compute and compare every hash against the base build:
python # hashes.py - compare identity and similarity hashes across variants import sys, hashlib, pefile, ppdeep, tlsh def row(path): data = open(path, "rb").read() pe = pefile.PE(data=data) return {"file": path, "sha256": hashlib.sha256(data).hexdigest(), "imphash": pe.get_imphash(), "ssdeep": ppdeep.hash(data), "tlsh": tlsh.hash(data)} rows = [row(p) for p in sys.argv[1:]] base = rows[0] print(f"{'vs ' + base['file']:<18}{'sha256':>8}{'imphash':>9}{'ssdeep':>8}{'tlsh':>6}") for r in rows[1:]: same = lambda k: "same" if r[k] == base[k] else "diff" print(f"{r['file']:<18}{same('sha256'):>8}{same('imphash'):>9}" f"{ppdeep.compare(base['ssdeep'], r['ssdeep']):>8}" f"{tlsh.diff(base['tlsh'], r['tlsh']):>6}")bash python3 hashes.py v1_base.exe v1_again.exe v2_string.exe v5_flipped.exe \ v3_O2.exe v6_import.exe v4_symbols.exeA run with mingw-w64 GCC 15 produced the table below. Your numbers will differ with the compiler version, and the TLSH distances move by a few points between otherwise identical builds, but the pattern should hold. ssdeep is a similarity score (higher is closer); TLSH is a distance (lower is closer).
text vs v1_base.exe sha256 imphash ssdeep tlsh v1_again.exe diff same 100 1 v2_string.exe diff same 99 3 v5_flipped.exe diff same 100 1 v3_O2.exe diff same 65 32 v6_import.exe diff diff 71 48 v4_symbols.exe diff same 0 511 -
Explain
v1_again.exe. Same source, same compiler, yet a different SHA-256. Find the bytes that differ:bash cmp -l v1_base.exe v1_again.exeThe differences sit in the COFF
TimeDateStampand the optional headerCheckSum(see PE Headers). Every build is a new file hash, even with no code change at all. -
Check the Rich header hash. Run
python3 -c "import pefile; print(repr(pefile.PE('v1_base.exe').get_rich_header_hash()))". A mingw build returns an empty string. If you have MSVC, buildtriage.cwithcltwice from different Visual Studio versions and compare. -
Identify every variant with
file, and if you have a Windows VM, drop them into Detect It Easy and TrID. Rename one to.pdfand confirm that none of the tools care. -
Start your triage log with one row per variant.
Questions to answer: Why does v4_symbols.exe score 0 in ssdeep but still
get a finite TLSH distance? Which hash would still link v6_import.exe to the
base if the author also packed it with UPX? Which two hashes would you publish
as IOCs for this "family", and which would you keep internal as clustering
pivots? What would happen to all five hashes if the builder appended 4 KB of
random bytes to the overlay?
Key takeaways
- Identify files by magic bytes and structure, never by name or extension;
file, TrID and DIE are heuristics you confirm with a hex view. - SHA-256 names one exact file. A rebuild, a timestamp or a single flipped byte gives a new hash, so file hashes are precise but brittle IOCs.
- Imphash, ssdeep, TLSH and the Rich header hash each survive different changes. Use two or more together to cluster, and packing defeats most of them.
- Look up hashes, do not upload files: uploads can leak victim data and warn the attacker.
- Log every file with its hashes, type, source, parent and lookup date before you go deeper.