Lesson 11.5 · Automated & Advanced Analysis· 50 min
Building an Analysis Pipeline
Turn manual triage into a repeatable pipeline: deduplication, file-type routing, static enrichment, scoring, safe isolation and reproducible reports.
Objectives
- Decide which triage steps to automate and which judgements must stay with an analyst
- Design the stages of a pipeline from intake and deduplication to scoring, storage and reporting
- Build a pipeline that never executes samples and survives malformed files, hangs and crashes
- Record rule and tool versions so every result can be reproduced and compared later
- Measure false positives and coverage, and feed analyst corrections back into rules
In Module 3 you triaged one file at a time: hash it, identify it, read its strings and imports, check for packing, write a YARA rule, and write it all up in a triage report. That workflow is correct, and for one sample it is fast. For five hundred samples a day it is impossible. A SOC mail gateway, an EDR quarantine queue or a research feed produces files far faster than anyone can open PE-bear.
A pipeline runs the mechanical part of triage automatically, the same way every time, and hands analysts a ranked queue with the evidence already collected. This lesson covers what to automate, how to structure the stages, how to stop a pipeline from becoming a liability, and how to keep its results trustworthy over time. The lab builds a small but real one in Python.
Why automate, and what stays manual
Three pressures push teams towards automation:
- Volume. Most files that reach a SOC are duplicates, benign, or commodity malware already known. A pipeline disposes of those so humans see the rest.
- Consistency. Two analysts triaging the same file will check different things on a busy day. A pipeline runs every check on every file, and its output always has the same fields.
- Time to verdict. An incident responder waiting on "is this malicious?" needs the hash reputation, packer status and rule hits in seconds, not after the analyst finishes lunch.
Automation does not replace judgement. It replaces lookups and bookkeeping.
| Automate | Keep manual |
|---|---|
| Hashing, deduplication, reputation lookup by hash | Final verdict on anything not already known |
| File-type identification and routing | Deciding whether a sample deserves deep reversing |
| Section entropy, imports, imphash, fuzzy hashes | Interpreting why an import combination is present |
| Running YARA and capa rule sets | Writing and promoting new detection rules |
| Sandbox detonation of routine samples | Handling targeted samples, or ones from a sensitive incident |
| Producing a structured report and IOC export | Attribution, confidence statements, and anything sent outside the team |
The line moves over time: a question analysts answer the same way ten times in a row is a candidate for a rule. But a pipeline produces evidence and a priority, not a conclusion, and its output should say so.
Pipeline stages
A practical pipeline is a series of stages, each consuming a sample plus everything earlier stages learned about it:
| Stage | Question it answers | Typical tools | Output |
|---|---|---|---|
| Intake and deduplication | Have we seen these exact bytes before? | SHA-256, sample store | New sample ID or pointer to existing result |
| File-type routing | What is it really, and which analysers apply? | Magic bytes, file/libmagic, TrID, Detect It Easy | Type tag that selects later stages |
| Static enrichment | What does the file look like without running it? | pefile, imphash, ssdeep/TLSH, FLOSS, capa, YARA, entropy | Structured facts |
| Dynamic detonation (optional) | What does it do when run? | CAPE, other sandboxes | Behaviour, dropped files, network, memory dumps |
| Scoring | How urgently should a human look? | Your own rules | Score, verdict band, reasons |
| Storage and indexing | Can we find it again, and related samples? | Object store, database, search index | Queryable records |
| Reporting and export | Who needs to know, in what format? | JSON, STIX 2.1, MISP | Report, IOCs, tickets |
Intake and deduplication
Every sample is identified by its SHA-256 as soon as it arrives, before any parser touches it. If the hash is already in the store, the pipeline records the new name, source and time against the existing sample and stops: running the same file through the same rules again gives the same answer. Keep all the names and sources, though. "This hash arrived from three different hosts under three different names" is itself intelligence, and it is the first thing an incident responder asks.
Hash lookups against external reputation services also belong here, and only the hash leaves your network (see the data-sharing warning below).
File-type routing
Route on content, never on the extension. An attachment called invoice.pdf
that begins with MZ is a PE file, and the mismatch is itself worth a point
in the score. The magic numbers from
Identifying and Hashing Files
decide which analysers run: PE and ELF
files go to the executable analysers, OLE and ZIP-based Office files to the
document tools from Malicious Documents and Archives,
scripts to deobfuscators, archives to an extractor
that feeds each member back into intake as a child sample.
Unknown types should still get hashes, strings and YARA. An analyser that does not recognise a file is not evidence that the file is harmless.
Static enrichment
This is Module 3 as code. For a PE file:
- cryptographic hashes plus imphash, and fuzzy hashes (ssdeep, TLSH) for finding near-duplicates;
- per-section size and entropy, and packer indicators from Detecting Packers and Entropy;
- imports grouped as in Reading Capabilities from Imports,
and capa's rule matches with ATT&CK IDs (
capa -jemits JSON); - strings and decoded strings from FLOSS, filtered for URLs, IPs, paths and mutex-like names;
- YARA matches from your team's rule set, as written in Writing Your First YARA Rules.
Each analyser writes into its own section of the result. If capa fails, the imports and YARA results must still be there.
Optional sandbox detonation
Detonation answers questions static analysis cannot, especially for packed samples whose code is invisible on disk. It is also expensive, slow (minutes per sample) and fooled by sandbox evasion and sleep tricks, which the planned lesson "Anti-VM and Sandbox Evasion" covers. Most pipelines detonate selectively: new hashes only, executable types only, and only when static scoring leaves the verdict open. Behaviour, dropped files and memory dumps come back as new evidence, and dropped files re-enter intake as children.
Scoring
The scorer turns evidence into a priority. Keep it additive and explainable: each rule contributes points and a sentence saying why, and the report lists those sentences. A score an analyst cannot explain to an incident lead is a score nobody will trust. Map score ranges to a few bands (for example low, review, suspicious, known-malicious) rather than pretending to a precise probability.
Two design rules matter more than the weights. A strong family-specific match
(a tested MAL_ YARA rule, a known-bad hash) should decide the band on its
own; generic traits such as "packed" or "few imports" should only raise
priority. And a stage that failed must never look like a clean result. If
the PE parser crashed, the sample is "incomplete", not "low".
Storage, reporting and export
Store the original sample once, by hash, in a write-once object store with restricted access. Store results separately, keyed by the same hash and by the pipeline run that produced them, so re-analysis with newer rules adds a result instead of overwriting history. Reports then come in two forms: a human-readable summary for the analyst queue, and machine-readable exports for other systems, covered below.
Designing for safety
A pipeline host processes hostile input all day, unattended. Treat it as exposed infrastructure, not as a convenient analyst workstation.
- Never execute samples on the pipeline host. Static analysers read bytes; detonation happens only in a separate, disposable sandbox network like the one in Building a Safe Analysis Lab. The pipeline host sends files to the sandbox and receives results; it never double-clicks anything.
- Isolate the analysers. Run each analyser for each file in its own process or container, as an unprivileged user, with no outbound network except to the services it needs. Parsers have bugs, and hostile files are exactly the input that triggers them.
- Expect malformed files. Truncated downloads, corrupt carves and deliberately broken headers are normal input. Catch exceptions per analyser, record them in the result, and continue.
- Bound every resource. Set a per-file wall-clock timeout, a maximum file
size, CPU and memory limits (cgroups, container limits, or
setrlimiton Linux), and a maximum number of strings or YARA matches recorded. - Handle archives defensively. Password-protected ZIP and 7z files are
how samples are exchanged (the password
infectedis a long-standing convention) and how attackers bypass mail scanners. Try known passwords, extract inside the worker, and cap the total extracted size, the compression ratio and the nesting depth so a decompression bomb exhausts a limit rather than the disk. - Store samples inert. Keep them in a store that is not an executable path, never on a share users can browse, and serve them to analysts only inside password-protected archives.
Warning: Treat everything in a sample-derived report as untrusted text. Strings extracted from malware end up in web dashboards, tickets and chat messages; escape them for HTML, defang URLs, and never let a pipeline follow a URL it found in a sample.
Data model and output formats
The core record is one JSON document per sample per analysis run. It should contain the identity of the sample (hashes, size, type, every name and source seen), each analyser's output in its own key, errors, the score with its reasons, and provenance: pipeline version, rule-set version and tool versions. The lab produces exactly this shape.
From that record you derive exports:
- STIX 2.1 for exchange with other organisations and TIPs. A sample
becomes a
fileobject with its hashes; a conclusion becomes anindicatorwith a pattern such as[file:hashes.'SHA-256' = '...'], linked by relationships to amalwareobject when a family is known. - MISP events, whose attributes (hashes, filename, imphash, ssdeep, TLSH) and objects map directly onto the enrichment fields. PyMISP makes this a few lines of code, and MISP's correlation engine then links your sample to other events that share any attribute.
- A search index (Elasticsearch, OpenSearch, or a relational database with good indexes) over the fields analysts pivot on: imphash, TLSH, YARA rule names, capa rule names, section names, PDB paths, extracted domains.
Index fields for pivoting, not for display. "Show me every sample whose imphash matches this one, or whose TLSH distance to it is under 50" is the query that turns a single triage into a cluster.
Existing platforms and where they fit
You rarely need to build all of this from scratch. The open-source ecosystem covers most stages; your own code is usually the glue and the scoring.
| Platform | Maintainer | What it is | Where it fits |
|---|---|---|---|
| CAPE Sandbox (CAPEv2) | Open-source community (kevoreilly) | Cuckoo-derived sandbox focused on unpacking, payload dumping and malware configuration extraction | The detonation stage, and a source of dumped payloads and configs |
| Assemblyline | Canadian Centre for Cyber Security | Scalable file triage platform: type identification, dozens of analysis services, scoring, and a web UI | A complete pipeline for teams that want one product rather than components |
| Karton | CERT Polska | Distributed framework in which small services consume and produce tasks, routed by headers such as file type | The orchestration layer for a pipeline built from your own services |
| MWDB Core | CERT Polska | Malware repository storing samples, extracted configurations and blobs with relations | Storage and search; pairs naturally with Karton |
| MISP | MISP Project | Threat-intelligence sharing platform with events, attributes, correlation and STIX support | The sharing and correlation stage |
| VirusTotal-style services | Commercial | Multi-engine scanning, sandboxes and retrohunting over a shared corpus | Reputation lookup and hunting, with caveats |
Warning: Uploading a file to a public multi-scanner shares it, often with every paying subscriber who can download it. That can leak a customer document, a phishing lure with a victim's name in it, or tell an attacker that their targeted implant has been found. Automate hash lookups, not uploads, and make uploading a manual decision governed by the incident's TLP.
Reproducibility: versioning rules and tools
A result is only meaningful if you can say what produced it. When a YARA rule is edited, a capa rule set updated or pefile upgraded, the same sample can score differently, and an analyst comparing last month's report with today's must be able to tell whether the sample changed or the pipeline did.
- Keep rules in Git, with a version in each rule's metadata, and record the commit (or a hash of the compiled rule file) in every result.
- Pin analyser versions (a lockfile, a container image digest) and record them in every result.
- Version the pipeline itself and its scoring weights. A change in threshold is a change in output.
- Never overwrite results. Re-analysis creates a new result linked to the old one, so you can diff them.
The lab writes a provenance block into every report for exactly this reason.
Measuring quality and closing the loop
A pipeline that nobody measures drifts. Rules are added and never removed, noisy heuristics creep up, and analysts learn to ignore the score. Track at least:
| Measure | How to get it | What it tells you |
|---|---|---|
| False-positive rate per rule | Run every rule change against a goodware corpus (clean OS installs, common software) before release, and count analyst "benign" overrides in production | Which rules to tighten or demote to hunting |
| Coverage | Run a labelled set of known-malicious samples through the pipeline; count how many reach "suspicious" or better | Which families or file types you miss |
| Error and timeout rate per analyser | Count errors entries by stage | Which parsers need fixing or limits adjusting |
| Time to verdict | Intake timestamp to first analyst decision | Whether the pipeline is actually saving time |
The feedback loop is what makes these numbers improve. When an analyst overrides a verdict, the override should be captured as data (sample hash, old verdict, new verdict, reason) and reviewed weekly. A false positive becomes a goodware test case and a tightened rule; a missed sample becomes a true-positive test case and a new rule. Both go through the same versioned, tested release as any other rule change.
Tip: Keep every sample that ever caused a false positive or a miss in a regression corpus, and run it in CI on each rule change. That corpus is worth more than any single rule in it.
Lab: a folder-in, JSON-out triage pipeline
You will build triage.py, which takes a folder, deduplicates by SHA-256,
identifies file types from magic bytes, enriches PE files (sections with
entropy, imports, imphash), computes ssdeep and TLSH, runs a small YARA rule
set, scores each sample with stated reasons, and writes one JSON report per
unique sample plus a summary table. Every input is a harmless program you
compile yourself. The pipeline never executes anything it analyses.
You need mingw-w64, UPX and Python 3. Everything else goes in a virtual
environment.
-
Create a working directory and a virtual environment with the analysis libraries:
bash mkdir -p m11p/src m11p/inbox m11p/rules && cd m11p python3 -m venv .venv .venv/bin/pip install pefile yara-python ppdeep python-tlshppdeepis a pure-Python ssdeep implementation andpython-tlshwraps TLSH, so no system libraries are required. You can addflare-capalater as another analyser. -
Create two benign programs.
src/hello.c:c #include <stdio.h> int main(void) { puts("hello from the lab"); return 0; }src/lab_agent.ccontains an implant-like user-agent string and looks up one harmless API by name, but does nothing else:c /* lab_agent.c - harmless: prints an implant-like string and resolves one API by name. */ #include <windows.h> #include <stdio.h> static const char USER_AGENT[] = "LabAgent/1.0 (Windows NT 10.0; lab-build)"; int main(void) { HMODULE k32 = GetModuleHandleA("kernel32.dll"); FARPROC p = GetProcAddress(k32, "GetTickCount"); printf("%s %p\n", USER_AGENT, (void *)p); return 0; } -
Fill the inbox with a realistic mix: two programs, a UPX-packed copy, an exact duplicate under another name, a text file, and a PE truncated to its first 400 bytes, the kind of broken file a failed download produces.
bash x86_64-w64-mingw32-gcc -O2 -s -o inbox/hello.exe src/hello.c x86_64-w64-mingw32-gcc -O2 -s -o inbox/lab_agent.exe src/lab_agent.c upx -q -o inbox/hello_upx.exe inbox/hello.exe cp inbox/hello.exe inbox/copy_of_hello.exe printf 'Meeting notes: rotate the lab VM snapshot on Friday.\n' > inbox/notes.txt head -c 400 inbox/lab_agent.exe > inbox/truncated.exe -
Write the rule set,
rules/triage.yar. Each rule carries ascorein its metadata, so the weights live next to the rule and are versioned with it:text import "pe" rule SUSP_PE_UPX_Sections { meta: description = "PE with UPX section names (packed; contents hidden from static analysis)" score = 30 condition: uint16(0) == 0x5A4D and for any s in pe.sections : ( s.name startswith "UPX" ) } rule SUSP_PE_Dynamic_API_Resolution { meta: description = "Imports GetProcAddress; may resolve APIs at runtime (common in benign code too)" score = 10 condition: uint16(0) == 0x5A4D and pe.imports("kernel32.dll", "GetProcAddress") } rule LAB_LabAgent_UserAgent { meta: description = "Lab training string LabAgent/x.y (stand-in for a family-specific artefact)" score = 40 strings: $ua = /LabAgent\/[0-9]{1,2}\.[0-9]{1,2}/ ascii wide condition: uint16(0) == 0x5A4D and $ua }LAB_LabAgent_UserAgentstands in for a tested family rule. The other two are generic traits, deliberately worth less. -
Write
triage.py:python #!/usr/bin/env python3 """triage.py - a minimal static triage pipeline for a folder of samples. Never executes samples. Each file is analysed in a separate worker process with a timeout, so a hung or crashing parser cannot stop the batch. """ import argparse, hashlib, json, math, multiprocessing as mp, os, sys, time from multiprocessing.connection import wait from datetime import datetime, timezone PIPELINE_VERSION = "0.3.0" MAX_SIZE = 50 * 1024 * 1024 # refuse files above 50 MB SUSPICIOUS_IMPORTS = {"VirtualAllocEx", "WriteProcessMemory", "CreateRemoteThread", "SetWindowsHookExA", "SetWindowsHookExW", "URLDownloadToFileA", "InternetOpenA", "InternetOpenW", "WinHttpOpen"} def entropy(data: bytes) -> float: if not data: return 0.0 counts = [0] * 256 for b in data: counts[b] += 1 n = len(data) return -sum(c / n * math.log2(c / n) for c in counts if c) def identify(head: bytes) -> str: """File type from magic bytes, never from the extension.""" if head[:2] == b"MZ": return "pe" if head[:4] == b"\x7fELF": return "elf" if head[:4] == b"PK\x03\x04": return "zip" if head[:4] == b"%PDF": return "pdf" if head[:8] == bytes.fromhex("d0cf11e0a1b11ae1"): return "ole" try: head.decode("utf-8") return "text" except UnicodeDecodeError: return "unknown" def pe_enrich(data: bytes) -> dict: import pefile pe = pefile.PE(data=data) sections = [{ "name": s.Name.rstrip(b"\x00").decode(errors="replace"), "raw_size": s.SizeOfRawData, "virtual_size": s.Misc_VirtualSize, "entropy": round(s.get_entropy(), 2), } for s in pe.sections] imports = {} for entry in getattr(pe, "DIRECTORY_ENTRY_IMPORT", []): dll = entry.dll.decode(errors="replace") imports[dll] = sorted(i.name.decode(errors="replace") for i in entry.imports if i.name) return { "machine": hex(pe.FILE_HEADER.Machine), "timestamp": pe.FILE_HEADER.TimeDateStamp, "imphash": pe.get_imphash() or None, "sections": sections, "imports": imports, "import_count": sum(len(v) for v in imports.values()), "warnings": pe.get_warnings()[:5], } def yara_scan(rules_path: str, data: bytes) -> list: import yara rules = yara.compile(filepath=rules_path) return [{"rule": m.rule, "score": int(m.meta.get("score", 0)), "description": m.meta.get("description", "")} for m in rules.match(data=data, timeout=10)] def score(report: dict) -> tuple: """Transparent additive scoring. Every point has a stated reason.""" points, reasons = 0, [] for hit in report.get("yara", []): points += hit["score"] reasons.append(f"+{hit['score']} yara:{hit['rule']}") pe = report.get("pe") if pe: hot = [s["name"] for s in pe["sections"] if s["entropy"] >= 7.2] if hot: points += 20 reasons.append(f"+20 high-entropy sections {hot}") if pe["import_count"] < 10: points += 15 reasons.append(f"+15 only {pe['import_count']} imports") sus = sorted({f for fs in pe["imports"].values() for f in fs} & SUSPICIOUS_IMPORTS) if sus: points += 10 * len(sus) reasons.append(f"+{10 * len(sus)} suspicious imports {sus}") verdict = ("suspicious" if points >= 50 else "review" if points >= 25 else "low") if report.get("errors") and verdict == "low": verdict = "incomplete" # a failed stage must never read as "clean" return points, verdict, reasons def analyse(path: str, rules_path: str, out) -> None: """Runs in a worker process. Every stage catches its own errors.""" report = {"errors": []} with open(path, "rb") as f: data = f.read(MAX_SIZE + 1) if len(data) > MAX_SIZE: out.send({"errors": ["file exceeds size limit"], "verdict": "error"}) return report["file_type"] = identify(data[:64]) report["size"] = len(data) report["entropy"] = round(entropy(data), 2) import ppdeep, tlsh report["ssdeep"] = ppdeep.hash(data) report["tlsh"] = tlsh.hash(data) if len(data) >= 50 else None if report["file_type"] == "pe": try: report["pe"] = pe_enrich(data) except Exception as e: # malformed PE: record, carry on report["errors"].append(f"pe: {type(e).__name__}: {e}") try: report["yara"] = yara_scan(rules_path, data) except Exception as e: report["errors"].append(f"yara: {type(e).__name__}: {e}") report["score"], report["verdict"], report["reasons"] = score(report) out.send(report) def run_with_timeout(path: str, rules_path: str, timeout: float) -> dict: recv, send = mp.Pipe(duplex=False) p = mp.Process(target=analyse, args=(path, rules_path, send)) p.start() send.close() # parent keeps only the read end if not wait([recv, p.sentinel], timeout): p.kill() p.join() return {"errors": [f"timeout after {timeout}s"], "verdict": "error"} try: result = recv.recv() except EOFError: # worker died without reporting p.join() return {"errors": [f"worker exited with code {p.exitcode}"], "verdict": "error"} p.join() return result def tool_versions() -> dict: import pefile, yara, ppdeep, tlsh from importlib.metadata import version return {"python": sys.version.split()[0], "pefile": pefile.__version__, "yara": yara.__version__, "ppdeep": version("ppdeep"), "python-tlsh": version("python-tlsh")} def main() -> None: ap = argparse.ArgumentParser() ap.add_argument("inbox") ap.add_argument("--rules", default="rules/triage.yar") ap.add_argument("--out", default="reports") ap.add_argument("--timeout", type=float, default=30.0) args = ap.parse_args() os.makedirs(args.out, exist_ok=True) with open(args.rules, "rb") as f: rules_sha256 = hashlib.sha256(f.read()).hexdigest() provenance = {"pipeline": PIPELINE_VERSION, "rules_sha256": rules_sha256, "tools": tool_versions()} # Stage 1: intake and deduplication by SHA-256. seen = {} for name in sorted(os.listdir(args.inbox)): path = os.path.join(args.inbox, name) if not os.path.isfile(path): continue h = hashlib.sha256() with open(path, "rb") as f: for chunk in iter(lambda: f.read(1 << 20), b""): h.update(chunk) seen.setdefault(h.hexdigest(), []).append(name) total = sum(len(v) for v in seen.values()) print(f"intake: {total} files, {len(seen)} unique, {total - len(seen)} duplicate(s)") # Stages 2-5: route, enrich, scan, score, store. rows = [] for sha256, names in seen.items(): start = time.monotonic() result = run_with_timeout(os.path.join(args.inbox, names[0]), args.rules, args.timeout) report = {"sha256": sha256, "names": names, "analysed_at": datetime.now(timezone.utc).isoformat(timespec="seconds"), "elapsed_s": round(time.monotonic() - start, 2), "provenance": provenance, **result} with open(os.path.join(args.out, f"{sha256}.json"), "w") as f: json.dump(report, f, indent=2) rows.append(report) print(f"{'sha256':<12} {'type':<7} {'score':>5} {'verdict':<10} {'names':<34} notes") for r in sorted(rows, key=lambda r: -r.get("score", -1)): notes = "; ".join(r.get("errors", [])) or ", ".join(h["rule"] for h in r.get("yara", [])) print(f"{r['sha256'][:12]} {r.get('file_type', '?'):<7} {r.get('score', '-'):>5} " f"{r.get('verdict'):<10} {','.join(r['names']):<34} {notes}") if __name__ == "__main__": main()Points to notice. Each unique file is analysed in a separate process; the parent waits on the result pipe and the process sentinel at once, and kills the worker when the timeout expires. A worker that crashes outright closes the pipe, which the parent reports as an error instead of hanging. The PE and YARA stages each catch their own exceptions. The
if __name__ == "__main__"guard is required, because on macOS and Windowsmultiprocessingstarts workers by re-importing the script. -
Run it:
bash .venv/bin/python triage.py inboxtext intake: 6 files, 5 unique, 1 duplicate(s) sha256 type score verdict names notes 988c0321ddbe pe 60 suspicious hello_upx.exe SUSP_PE_UPX_Sections, SUSP_PE_Dynamic_API_Resolution ac62887ad766 pe 50 suspicious lab_agent.exe SUSP_PE_Dynamic_API_Resolution, LAB_LabAgent_UserAgent 78afa14aa60b pe 0 low copy_of_hello.exe,hello.exe 7616f0c28322 text 0 low notes.txt 0d5f29750958 pe 0 incomplete truncated.exe pe: PEFormatError: 'Data length less than expected header length.'Your hashes will differ; the shape will not. The duplicate was analysed once and both names were kept. The truncated PE made pefile raise
PEFormatError, the stage recorded it, YARA and the fuzzy hashes still ran, and the verdict isincompleterather thanlow. And the benign UPX-packedhellooutranks everything: generic packer traits added up to 60 points. That is the false positive you will be asked about below. -
Open a report. Trimmed, the one for the packed file reads:
json { "sha256": "988c0321ddbe2f9933722bbcaa4dfe82250e48cdb36d5f93655d8259f973776e", "names": ["hello_upx.exe"], "provenance": { "pipeline": "0.3.0", "rules_sha256": "e40108ed7fa7ba8a2afab1e2ce6ab9936104da82b95c33556a3303c88d094afd", "tools": { "python": "3.14.7", "pefile": "2024.8.26", "yara": "4.5.4", "ppdeep": "20260221", "python-tlsh": "4.5.0" } }, "errors": [], "file_type": "pe", "size": 8704, "entropy": 6.99, "ssdeep": "192:6E842x44/bJqkxZLrgJqfsPLFr7oyzBLgPWk3S2KD:6lAmJX3D6r7okBIzK", "tlsh": "T19D029E9B3469015BD6140EBFB1E21D58ACA17C17FB57A324CFB000A215866BB58BEF1F", "pe": { "imphash": "fbb1ab502fa1db2c50c71180a9299c72", "sections": [ { "name": "UPX0", "raw_size": 0, "virtual_size": 40960, "entropy": 0.0 }, { "name": "UPX1", "raw_size": 7168, "virtual_size": 8192, "entropy": 7.48 }, { "name": "UPX2", "raw_size": 1024, "virtual_size": 4096, "entropy": 3.31 } ], "imports": { "KERNEL32.DLL": ["ExitProcess", "GetProcAddress", "LoadLibraryA", "VirtualProtect"] }, "import_count": 12 }, "score": 60, "verdict": "suspicious", "reasons": [ "+30 yara:SUSP_PE_UPX_Sections", "+10 yara:SUSP_PE_Dynamic_API_Resolution", "+20 high-entropy sections ['UPX1']" ] }The eight single-import CRT DLLs are trimmed from
imports. Note thatGetProcAddresshere comes from the UPX stub, which rebuilds the real import table at runtime, not fromhello.c. -
Exercise the timeout path. You do not need a hostile file to see it: set a timeout shorter than a worker takes to start.
bash .venv/bin/python triage.py inbox --timeout 0.01 --out reports_ttext intake: 6 files, 5 unique, 1 duplicate(s) sha256 type score verdict names notes 78afa14aa60b ? - error copy_of_hello.exe,hello.exe timeout after 0.01s 988c0321ddbe ? - error hello_upx.exe timeout after 0.01s ac62887ad766 ? - error lab_agent.exe timeout after 0.01s 7616f0c28322 ? - error notes.txt timeout after 0.01s 0d5f29750958 ? - error truncated.exe timeout after 0.01sEvery worker was killed and every sample still got a report saying so. The batch finished.
-
Use the fuzzy hashes. Compare
hello.exewithlab_agent.exeand with the packed copy:bash .venv/bin/python - <<'EOF' import json, glob, ppdeep, tlsh r = {json.load(open(f))["names"][-1]: json.load(open(f)) for f in glob.glob("reports/*.json")} a, b, c = r["hello.exe"], r["lab_agent.exe"], r["hello_upx.exe"] print("ssdeep hello/lab_agent", ppdeep.compare(a["ssdeep"], b["ssdeep"]), "hello/upx", ppdeep.compare(a["ssdeep"], c["ssdeep"])) print("tlsh hello/lab_agent", tlsh.diff(a["tlsh"], b["tlsh"]), "hello/upx", tlsh.diff(a["tlsh"], c["tlsh"])) EOFtext ssdeep hello/lab_agent 0 hello/upx 0 tlsh hello/lab_agent 36 hello/upx 332ssdeep (a similarity score from 0 to 100) finds no resemblance at all. TLSH (a distance, where lower is closer) places the two small mingw-w64 programs close together, because they share most of their runtime code, and the packed copy far away, because packing changes every byte. Neither hash sees through packing. That is why packed samples go to unpacking or detonation before similarity search.
Questions to answer: Which scoring change would stop the benign packed
hello_upx.exe from being "suspicious" without hiding a genuinely packed
malicious sample, and how would you test that change? What would happen to
the batch if analyse ran in the parent process and pefile entered an
infinite loop on one file? Why does the pipeline record rules_sha256, and
what question does it let you answer six months from now? How would you
extend intake to handle a password-protected ZIP without ever writing its
members to a directory users can reach? Which fields of these reports would
you index first, and which pivot query would you run on lab_agent.exe?
Key takeaways
- A pipeline automates lookups and bookkeeping (hashing, routing, enrichment, rule scanning, reporting) and hands analysts a ranked queue; verdicts on new samples, rule promotion and external sharing stay with people.
- Deduplicate by SHA-256 at intake, route on magic bytes rather than extensions, and keep every analyser's output, and its errors, separate.
- Score additively with a stated reason for every point, let tested family rules decide on their own, and never let a failed stage read as a clean one.
- Never execute samples on the pipeline host; isolate each analyser, bound its time, memory and output, and treat archives and extracted strings as hostile.
- Store one JSON record per sample per run with provenance (pipeline, rule and tool versions), and derive STIX, MISP and search-index views from it.
- Reuse existing platforms (CAPE, Assemblyline, Karton and MWDB, MISP) where they fit, look up hashes rather than uploading files, and measure false positives and coverage so analyst feedback turns into better rules.