Skip to content

Leçon 11.5 · Analyse automatisée & avancée· 50 min

Building an Analysis Pipeline

Turn manual triage into a repeatable pipeline: deduplication, file-type routing, static enrichment, scoring, safe isolation and reproducible reports.

Cette leçon n’est disponible qu’en anglais pour le moment.

Objectifs

  • Decide which triage steps to automate and which judgements must stay with an analyst
  • Design the stages of a pipeline from intake and deduplication to scoring, storage and reporting
  • Build a pipeline that never executes samples and survives malformed files, hangs and crashes
  • Record rule and tool versions so every result can be reproduced and compared later
  • Measure false positives and coverage, and feed analyst corrections back into rules

In Module 3 you triaged one file at a time: hash it, identify it, read its strings and imports, check for packing, write a YARA rule, and write it all up in a triage report. That workflow is correct, and for one sample it is fast. For five hundred samples a day it is impossible. A SOC mail gateway, an EDR quarantine queue or a research feed produces files far faster than anyone can open PE-bear.

A pipeline runs the mechanical part of triage automatically, the same way every time, and hands analysts a ranked queue with the evidence already collected. This lesson covers what to automate, how to structure the stages, how to stop a pipeline from becoming a liability, and how to keep its results trustworthy over time. The lab builds a small but real one in Python.

Why automate, and what stays manual

Three pressures push teams towards automation:

  • Volume. Most files that reach a SOC are duplicates, benign, or commodity malware already known. A pipeline disposes of those so humans see the rest.
  • Consistency. Two analysts triaging the same file will check different things on a busy day. A pipeline runs every check on every file, and its output always has the same fields.
  • Time to verdict. An incident responder waiting on "is this malicious?" needs the hash reputation, packer status and rule hits in seconds, not after the analyst finishes lunch.

Automation does not replace judgement. It replaces lookups and bookkeeping.

AutomateKeep manual
Hashing, deduplication, reputation lookup by hashFinal verdict on anything not already known
File-type identification and routingDeciding whether a sample deserves deep reversing
Section entropy, imports, imphash, fuzzy hashesInterpreting why an import combination is present
Running YARA and capa rule setsWriting and promoting new detection rules
Sandbox detonation of routine samplesHandling targeted samples, or ones from a sensitive incident
Producing a structured report and IOC exportAttribution, confidence statements, and anything sent outside the team

The line moves over time: a question analysts answer the same way ten times in a row is a candidate for a rule. But a pipeline produces evidence and a priority, not a conclusion, and its output should say so.

Pipeline stages

A practical pipeline is a series of stages, each consuming a sample plus everything earlier stages learned about it:

StageQuestion it answersTypical toolsOutput
Intake and deduplicationHave we seen these exact bytes before?SHA-256, sample storeNew sample ID or pointer to existing result
File-type routingWhat is it really, and which analysers apply?Magic bytes, file/libmagic, TrID, Detect It EasyType tag that selects later stages
Static enrichmentWhat does the file look like without running it?pefile, imphash, ssdeep/TLSH, FLOSS, capa, YARA, entropyStructured facts
Dynamic detonation (optional)What does it do when run?CAPE, other sandboxesBehaviour, dropped files, network, memory dumps
ScoringHow urgently should a human look?Your own rulesScore, verdict band, reasons
Storage and indexingCan we find it again, and related samples?Object store, database, search indexQueryable records
Reporting and exportWho needs to know, in what format?JSON, STIX 2.1, MISPReport, IOCs, tickets

Intake and deduplication

Every sample is identified by its SHA-256 as soon as it arrives, before any parser touches it. If the hash is already in the store, the pipeline records the new name, source and time against the existing sample and stops: running the same file through the same rules again gives the same answer. Keep all the names and sources, though. "This hash arrived from three different hosts under three different names" is itself intelligence, and it is the first thing an incident responder asks.

Hash lookups against external reputation services also belong here, and only the hash leaves your network (see the data-sharing warning below).

File-type routing

Route on content, never on the extension. An attachment called invoice.pdf that begins with MZ is a PE file, and the mismatch is itself worth a point in the score. The magic numbers from Identifying and Hashing Files decide which analysers run: PE and ELF files go to the executable analysers, OLE and ZIP-based Office files to the document tools from Malicious Documents and Archives, scripts to deobfuscators, archives to an extractor that feeds each member back into intake as a child sample.

Unknown types should still get hashes, strings and YARA. An analyser that does not recognise a file is not evidence that the file is harmless.

Static enrichment

This is Module 3 as code. For a PE file:

Each analyser writes into its own section of the result. If capa fails, the imports and YARA results must still be there.

Optional sandbox detonation

Detonation answers questions static analysis cannot, especially for packed samples whose code is invisible on disk. It is also expensive, slow (minutes per sample) and fooled by sandbox evasion and sleep tricks, which the planned lesson "Anti-VM and Sandbox Evasion" covers. Most pipelines detonate selectively: new hashes only, executable types only, and only when static scoring leaves the verdict open. Behaviour, dropped files and memory dumps come back as new evidence, and dropped files re-enter intake as children.

Scoring

The scorer turns evidence into a priority. Keep it additive and explainable: each rule contributes points and a sentence saying why, and the report lists those sentences. A score an analyst cannot explain to an incident lead is a score nobody will trust. Map score ranges to a few bands (for example low, review, suspicious, known-malicious) rather than pretending to a precise probability.

Two design rules matter more than the weights. A strong family-specific match (a tested MAL_ YARA rule, a known-bad hash) should decide the band on its own; generic traits such as "packed" or "few imports" should only raise priority. And a stage that failed must never look like a clean result. If the PE parser crashed, the sample is "incomplete", not "low".

Storage, reporting and export

Store the original sample once, by hash, in a write-once object store with restricted access. Store results separately, keyed by the same hash and by the pipeline run that produced them, so re-analysis with newer rules adds a result instead of overwriting history. Reports then come in two forms: a human-readable summary for the analyst queue, and machine-readable exports for other systems, covered below.

Designing for safety

A pipeline host processes hostile input all day, unattended. Treat it as exposed infrastructure, not as a convenient analyst workstation.

  • Never execute samples on the pipeline host. Static analysers read bytes; detonation happens only in a separate, disposable sandbox network like the one in Building a Safe Analysis Lab. The pipeline host sends files to the sandbox and receives results; it never double-clicks anything.
  • Isolate the analysers. Run each analyser for each file in its own process or container, as an unprivileged user, with no outbound network except to the services it needs. Parsers have bugs, and hostile files are exactly the input that triggers them.
  • Expect malformed files. Truncated downloads, corrupt carves and deliberately broken headers are normal input. Catch exceptions per analyser, record them in the result, and continue.
  • Bound every resource. Set a per-file wall-clock timeout, a maximum file size, CPU and memory limits (cgroups, container limits, or setrlimit on Linux), and a maximum number of strings or YARA matches recorded.
  • Handle archives defensively. Password-protected ZIP and 7z files are how samples are exchanged (the password infected is a long-standing convention) and how attackers bypass mail scanners. Try known passwords, extract inside the worker, and cap the total extracted size, the compression ratio and the nesting depth so a decompression bomb exhausts a limit rather than the disk.
  • Store samples inert. Keep them in a store that is not an executable path, never on a share users can browse, and serve them to analysts only inside password-protected archives.

Warning: Treat everything in a sample-derived report as untrusted text. Strings extracted from malware end up in web dashboards, tickets and chat messages; escape them for HTML, defang URLs, and never let a pipeline follow a URL it found in a sample.

Data model and output formats

The core record is one JSON document per sample per analysis run. It should contain the identity of the sample (hashes, size, type, every name and source seen), each analyser's output in its own key, errors, the score with its reasons, and provenance: pipeline version, rule-set version and tool versions. The lab produces exactly this shape.

From that record you derive exports:

  • STIX 2.1 for exchange with other organisations and TIPs. A sample becomes a file object with its hashes; a conclusion becomes an indicator with a pattern such as [file:hashes.'SHA-256' = '...'], linked by relationships to a malware object when a family is known.
  • MISP events, whose attributes (hashes, filename, imphash, ssdeep, TLSH) and objects map directly onto the enrichment fields. PyMISP makes this a few lines of code, and MISP's correlation engine then links your sample to other events that share any attribute.
  • A search index (Elasticsearch, OpenSearch, or a relational database with good indexes) over the fields analysts pivot on: imphash, TLSH, YARA rule names, capa rule names, section names, PDB paths, extracted domains.

Index fields for pivoting, not for display. "Show me every sample whose imphash matches this one, or whose TLSH distance to it is under 50" is the query that turns a single triage into a cluster.

Existing platforms and where they fit

You rarely need to build all of this from scratch. The open-source ecosystem covers most stages; your own code is usually the glue and the scoring.

PlatformMaintainerWhat it isWhere it fits
CAPE Sandbox (CAPEv2)Open-source community (kevoreilly)Cuckoo-derived sandbox focused on unpacking, payload dumping and malware configuration extractionThe detonation stage, and a source of dumped payloads and configs
AssemblylineCanadian Centre for Cyber SecurityScalable file triage platform: type identification, dozens of analysis services, scoring, and a web UIA complete pipeline for teams that want one product rather than components
KartonCERT PolskaDistributed framework in which small services consume and produce tasks, routed by headers such as file typeThe orchestration layer for a pipeline built from your own services
MWDB CoreCERT PolskaMalware repository storing samples, extracted configurations and blobs with relationsStorage and search; pairs naturally with Karton
MISPMISP ProjectThreat-intelligence sharing platform with events, attributes, correlation and STIX supportThe sharing and correlation stage
VirusTotal-style servicesCommercialMulti-engine scanning, sandboxes and retrohunting over a shared corpusReputation lookup and hunting, with caveats

Warning: Uploading a file to a public multi-scanner shares it, often with every paying subscriber who can download it. That can leak a customer document, a phishing lure with a victim's name in it, or tell an attacker that their targeted implant has been found. Automate hash lookups, not uploads, and make uploading a manual decision governed by the incident's TLP.

Reproducibility: versioning rules and tools

A result is only meaningful if you can say what produced it. When a YARA rule is edited, a capa rule set updated or pefile upgraded, the same sample can score differently, and an analyst comparing last month's report with today's must be able to tell whether the sample changed or the pipeline did.

  • Keep rules in Git, with a version in each rule's metadata, and record the commit (or a hash of the compiled rule file) in every result.
  • Pin analyser versions (a lockfile, a container image digest) and record them in every result.
  • Version the pipeline itself and its scoring weights. A change in threshold is a change in output.
  • Never overwrite results. Re-analysis creates a new result linked to the old one, so you can diff them.

The lab writes a provenance block into every report for exactly this reason.

Measuring quality and closing the loop

A pipeline that nobody measures drifts. Rules are added and never removed, noisy heuristics creep up, and analysts learn to ignore the score. Track at least:

MeasureHow to get itWhat it tells you
False-positive rate per ruleRun every rule change against a goodware corpus (clean OS installs, common software) before release, and count analyst "benign" overrides in productionWhich rules to tighten or demote to hunting
CoverageRun a labelled set of known-malicious samples through the pipeline; count how many reach "suspicious" or betterWhich families or file types you miss
Error and timeout rate per analyserCount errors entries by stageWhich parsers need fixing or limits adjusting
Time to verdictIntake timestamp to first analyst decisionWhether the pipeline is actually saving time

The feedback loop is what makes these numbers improve. When an analyst overrides a verdict, the override should be captured as data (sample hash, old verdict, new verdict, reason) and reviewed weekly. A false positive becomes a goodware test case and a tightened rule; a missed sample becomes a true-positive test case and a new rule. Both go through the same versioned, tested release as any other rule change.

Tip: Keep every sample that ever caused a false positive or a miss in a regression corpus, and run it in CI on each rule change. That corpus is worth more than any single rule in it.

Lab: a folder-in, JSON-out triage pipeline

You will build triage.py, which takes a folder, deduplicates by SHA-256, identifies file types from magic bytes, enriches PE files (sections with entropy, imports, imphash), computes ssdeep and TLSH, runs a small YARA rule set, scores each sample with stated reasons, and writes one JSON report per unique sample plus a summary table. Every input is a harmless program you compile yourself. The pipeline never executes anything it analyses.

You need mingw-w64, UPX and Python 3. Everything else goes in a virtual environment.

  1. Create a working directory and a virtual environment with the analysis libraries:

    bash
    mkdir -p m11p/src m11p/inbox m11p/rules && cd m11p
    python3 -m venv .venv
    .venv/bin/pip install pefile yara-python ppdeep python-tlsh

    ppdeep is a pure-Python ssdeep implementation and python-tlsh wraps TLSH, so no system libraries are required. You can add flare-capa later as another analyser.

  2. Create two benign programs. src/hello.c:

    c
    #include <stdio.h>
    int main(void) { puts("hello from the lab"); return 0; }

    src/lab_agent.c contains an implant-like user-agent string and looks up one harmless API by name, but does nothing else:

    c
    /* lab_agent.c - harmless: prints an implant-like string and resolves one API by name. */
    #include <windows.h>
    #include <stdio.h>
    static const char USER_AGENT[] = "LabAgent/1.0 (Windows NT 10.0; lab-build)";
    int main(void) {
        HMODULE k32 = GetModuleHandleA("kernel32.dll");
        FARPROC p = GetProcAddress(k32, "GetTickCount");
        printf("%s %p\n", USER_AGENT, (void *)p);
        return 0;
    }
  3. Fill the inbox with a realistic mix: two programs, a UPX-packed copy, an exact duplicate under another name, a text file, and a PE truncated to its first 400 bytes, the kind of broken file a failed download produces.

    bash
    x86_64-w64-mingw32-gcc -O2 -s -o inbox/hello.exe src/hello.c
    x86_64-w64-mingw32-gcc -O2 -s -o inbox/lab_agent.exe src/lab_agent.c
    upx -q -o inbox/hello_upx.exe inbox/hello.exe
    cp inbox/hello.exe inbox/copy_of_hello.exe
    printf 'Meeting notes: rotate the lab VM snapshot on Friday.\n' > inbox/notes.txt
    head -c 400 inbox/lab_agent.exe > inbox/truncated.exe
  4. Write the rule set, rules/triage.yar. Each rule carries a score in its metadata, so the weights live next to the rule and are versioned with it:

    text
    import "pe"
    
    rule SUSP_PE_UPX_Sections
    {
        meta:
            description = "PE with UPX section names (packed; contents hidden from static analysis)"
            score       = 30
        condition:
            uint16(0) == 0x5A4D
            and for any s in pe.sections : ( s.name startswith "UPX" )
    }
    
    rule SUSP_PE_Dynamic_API_Resolution
    {
        meta:
            description = "Imports GetProcAddress; may resolve APIs at runtime (common in benign code too)"
            score       = 10
        condition:
            uint16(0) == 0x5A4D and pe.imports("kernel32.dll", "GetProcAddress")
    }
    
    rule LAB_LabAgent_UserAgent
    {
        meta:
            description = "Lab training string LabAgent/x.y (stand-in for a family-specific artefact)"
            score       = 40
        strings:
            $ua = /LabAgent\/[0-9]{1,2}\.[0-9]{1,2}/ ascii wide
        condition:
            uint16(0) == 0x5A4D and $ua
    }

    LAB_LabAgent_UserAgent stands in for a tested family rule. The other two are generic traits, deliberately worth less.

  5. Write triage.py:

    python
    #!/usr/bin/env python3
    """triage.py - a minimal static triage pipeline for a folder of samples.
    
    Never executes samples. Each file is analysed in a separate worker process
    with a timeout, so a hung or crashing parser cannot stop the batch.
    """
    import argparse, hashlib, json, math, multiprocessing as mp, os, sys, time
    from multiprocessing.connection import wait
    from datetime import datetime, timezone
    
    PIPELINE_VERSION = "0.3.0"
    MAX_SIZE = 50 * 1024 * 1024            # refuse files above 50 MB
    SUSPICIOUS_IMPORTS = {"VirtualAllocEx", "WriteProcessMemory", "CreateRemoteThread",
                          "SetWindowsHookExA", "SetWindowsHookExW", "URLDownloadToFileA",
                          "InternetOpenA", "InternetOpenW", "WinHttpOpen"}
    
    def entropy(data: bytes) -> float:
        if not data:
            return 0.0
        counts = [0] * 256
        for b in data:
            counts[b] += 1
        n = len(data)
        return -sum(c / n * math.log2(c / n) for c in counts if c)
    
    def identify(head: bytes) -> str:
        """File type from magic bytes, never from the extension."""
        if head[:2] == b"MZ":
            return "pe"
        if head[:4] == b"\x7fELF":
            return "elf"
        if head[:4] == b"PK\x03\x04":
            return "zip"
        if head[:4] == b"%PDF":
            return "pdf"
        if head[:8] == bytes.fromhex("d0cf11e0a1b11ae1"):
            return "ole"
        try:
            head.decode("utf-8")
            return "text"
        except UnicodeDecodeError:
            return "unknown"
    
    def pe_enrich(data: bytes) -> dict:
        import pefile
        pe = pefile.PE(data=data)
        sections = [{
            "name": s.Name.rstrip(b"\x00").decode(errors="replace"),
            "raw_size": s.SizeOfRawData,
            "virtual_size": s.Misc_VirtualSize,
            "entropy": round(s.get_entropy(), 2),
        } for s in pe.sections]
        imports = {}
        for entry in getattr(pe, "DIRECTORY_ENTRY_IMPORT", []):
            dll = entry.dll.decode(errors="replace")
            imports[dll] = sorted(i.name.decode(errors="replace")
                                  for i in entry.imports if i.name)
        return {
            "machine": hex(pe.FILE_HEADER.Machine),
            "timestamp": pe.FILE_HEADER.TimeDateStamp,
            "imphash": pe.get_imphash() or None,
            "sections": sections,
            "imports": imports,
            "import_count": sum(len(v) for v in imports.values()),
            "warnings": pe.get_warnings()[:5],
        }
    
    def yara_scan(rules_path: str, data: bytes) -> list:
        import yara
        rules = yara.compile(filepath=rules_path)
        return [{"rule": m.rule, "score": int(m.meta.get("score", 0)),
                 "description": m.meta.get("description", "")}
                for m in rules.match(data=data, timeout=10)]
    
    def score(report: dict) -> tuple:
        """Transparent additive scoring. Every point has a stated reason."""
        points, reasons = 0, []
        for hit in report.get("yara", []):
            points += hit["score"]
            reasons.append(f"+{hit['score']} yara:{hit['rule']}")
        pe = report.get("pe")
        if pe:
            hot = [s["name"] for s in pe["sections"] if s["entropy"] >= 7.2]
            if hot:
                points += 20
                reasons.append(f"+20 high-entropy sections {hot}")
            if pe["import_count"] < 10:
                points += 15
                reasons.append(f"+15 only {pe['import_count']} imports")
            sus = sorted({f for fs in pe["imports"].values() for f in fs} & SUSPICIOUS_IMPORTS)
            if sus:
                points += 10 * len(sus)
                reasons.append(f"+{10 * len(sus)} suspicious imports {sus}")
        verdict = ("suspicious" if points >= 50 else "review" if points >= 25 else "low")
        if report.get("errors") and verdict == "low":
            verdict = "incomplete"     # a failed stage must never read as "clean"
        return points, verdict, reasons
    
    def analyse(path: str, rules_path: str, out) -> None:
        """Runs in a worker process. Every stage catches its own errors."""
        report = {"errors": []}
        with open(path, "rb") as f:
            data = f.read(MAX_SIZE + 1)
        if len(data) > MAX_SIZE:
            out.send({"errors": ["file exceeds size limit"], "verdict": "error"})
            return
        report["file_type"] = identify(data[:64])
        report["size"] = len(data)
        report["entropy"] = round(entropy(data), 2)
        import ppdeep, tlsh
        report["ssdeep"] = ppdeep.hash(data)
        report["tlsh"] = tlsh.hash(data) if len(data) >= 50 else None
        if report["file_type"] == "pe":
            try:
                report["pe"] = pe_enrich(data)
            except Exception as e:                    # malformed PE: record, carry on
                report["errors"].append(f"pe: {type(e).__name__}: {e}")
        try:
            report["yara"] = yara_scan(rules_path, data)
        except Exception as e:
            report["errors"].append(f"yara: {type(e).__name__}: {e}")
        report["score"], report["verdict"], report["reasons"] = score(report)
        out.send(report)
    
    def run_with_timeout(path: str, rules_path: str, timeout: float) -> dict:
        recv, send = mp.Pipe(duplex=False)
        p = mp.Process(target=analyse, args=(path, rules_path, send))
        p.start()
        send.close()                           # parent keeps only the read end
        if not wait([recv, p.sentinel], timeout):
            p.kill()
            p.join()
            return {"errors": [f"timeout after {timeout}s"], "verdict": "error"}
        try:
            result = recv.recv()
        except EOFError:                       # worker died without reporting
            p.join()
            return {"errors": [f"worker exited with code {p.exitcode}"], "verdict": "error"}
        p.join()
        return result
    
    def tool_versions() -> dict:
        import pefile, yara, ppdeep, tlsh
        from importlib.metadata import version
        return {"python": sys.version.split()[0], "pefile": pefile.__version__,
                "yara": yara.__version__, "ppdeep": version("ppdeep"),
                "python-tlsh": version("python-tlsh")}
    
    def main() -> None:
        ap = argparse.ArgumentParser()
        ap.add_argument("inbox")
        ap.add_argument("--rules", default="rules/triage.yar")
        ap.add_argument("--out", default="reports")
        ap.add_argument("--timeout", type=float, default=30.0)
        args = ap.parse_args()
        os.makedirs(args.out, exist_ok=True)
    
        with open(args.rules, "rb") as f:
            rules_sha256 = hashlib.sha256(f.read()).hexdigest()
        provenance = {"pipeline": PIPELINE_VERSION, "rules_sha256": rules_sha256,
                      "tools": tool_versions()}
    
        # Stage 1: intake and deduplication by SHA-256.
        seen = {}
        for name in sorted(os.listdir(args.inbox)):
            path = os.path.join(args.inbox, name)
            if not os.path.isfile(path):
                continue
            h = hashlib.sha256()
            with open(path, "rb") as f:
                for chunk in iter(lambda: f.read(1 << 20), b""):
                    h.update(chunk)
            seen.setdefault(h.hexdigest(), []).append(name)
        total = sum(len(v) for v in seen.values())
        print(f"intake: {total} files, {len(seen)} unique, {total - len(seen)} duplicate(s)")
    
        # Stages 2-5: route, enrich, scan, score, store.
        rows = []
        for sha256, names in seen.items():
            start = time.monotonic()
            result = run_with_timeout(os.path.join(args.inbox, names[0]), args.rules, args.timeout)
            report = {"sha256": sha256, "names": names,
                      "analysed_at": datetime.now(timezone.utc).isoformat(timespec="seconds"),
                      "elapsed_s": round(time.monotonic() - start, 2),
                      "provenance": provenance, **result}
            with open(os.path.join(args.out, f"{sha256}.json"), "w") as f:
                json.dump(report, f, indent=2)
            rows.append(report)
    
        print(f"{'sha256':<12} {'type':<7} {'score':>5}  {'verdict':<10} {'names':<34} notes")
        for r in sorted(rows, key=lambda r: -r.get("score", -1)):
            notes = "; ".join(r.get("errors", [])) or ", ".join(h["rule"] for h in r.get("yara", []))
            print(f"{r['sha256'][:12]} {r.get('file_type', '?'):<7} {r.get('score', '-'):>5}  "
                  f"{r.get('verdict'):<10} {','.join(r['names']):<34} {notes}")
    
    if __name__ == "__main__":
        main()

    Points to notice. Each unique file is analysed in a separate process; the parent waits on the result pipe and the process sentinel at once, and kills the worker when the timeout expires. A worker that crashes outright closes the pipe, which the parent reports as an error instead of hanging. The PE and YARA stages each catch their own exceptions. The if __name__ == "__main__" guard is required, because on macOS and Windows multiprocessing starts workers by re-importing the script.

  6. Run it:

    bash
    .venv/bin/python triage.py inbox
    text
    intake: 6 files, 5 unique, 1 duplicate(s)
    sha256       type    score  verdict    names                              notes
    988c0321ddbe pe         60  suspicious hello_upx.exe                      SUSP_PE_UPX_Sections, SUSP_PE_Dynamic_API_Resolution
    ac62887ad766 pe         50  suspicious lab_agent.exe                      SUSP_PE_Dynamic_API_Resolution, LAB_LabAgent_UserAgent
    78afa14aa60b pe          0  low        copy_of_hello.exe,hello.exe
    7616f0c28322 text        0  low        notes.txt
    0d5f29750958 pe          0  incomplete truncated.exe                      pe: PEFormatError: 'Data length less than expected header length.'

    Your hashes will differ; the shape will not. The duplicate was analysed once and both names were kept. The truncated PE made pefile raise PEFormatError, the stage recorded it, YARA and the fuzzy hashes still ran, and the verdict is incomplete rather than low. And the benign UPX-packed hello outranks everything: generic packer traits added up to 60 points. That is the false positive you will be asked about below.

  7. Open a report. Trimmed, the one for the packed file reads:

    json
    {
      "sha256": "988c0321ddbe2f9933722bbcaa4dfe82250e48cdb36d5f93655d8259f973776e",
      "names": ["hello_upx.exe"],
      "provenance": {
        "pipeline": "0.3.0",
        "rules_sha256": "e40108ed7fa7ba8a2afab1e2ce6ab9936104da82b95c33556a3303c88d094afd",
        "tools": { "python": "3.14.7", "pefile": "2024.8.26", "yara": "4.5.4",
                   "ppdeep": "20260221", "python-tlsh": "4.5.0" }
      },
      "errors": [],
      "file_type": "pe",
      "size": 8704,
      "entropy": 6.99,
      "ssdeep": "192:6E842x44/bJqkxZLrgJqfsPLFr7oyzBLgPWk3S2KD:6lAmJX3D6r7okBIzK",
      "tlsh": "T19D029E9B3469015BD6140EBFB1E21D58ACA17C17FB57A324CFB000A215866BB58BEF1F",
      "pe": {
        "imphash": "fbb1ab502fa1db2c50c71180a9299c72",
        "sections": [
          { "name": "UPX0", "raw_size": 0, "virtual_size": 40960, "entropy": 0.0 },
          { "name": "UPX1", "raw_size": 7168, "virtual_size": 8192, "entropy": 7.48 },
          { "name": "UPX2", "raw_size": 1024, "virtual_size": 4096, "entropy": 3.31 }
        ],
        "imports": {
          "KERNEL32.DLL": ["ExitProcess", "GetProcAddress", "LoadLibraryA", "VirtualProtect"]
        },
        "import_count": 12
      },
      "score": 60,
      "verdict": "suspicious",
      "reasons": [
        "+30 yara:SUSP_PE_UPX_Sections",
        "+10 yara:SUSP_PE_Dynamic_API_Resolution",
        "+20 high-entropy sections ['UPX1']"
      ]
    }

    The eight single-import CRT DLLs are trimmed from imports. Note that GetProcAddress here comes from the UPX stub, which rebuilds the real import table at runtime, not from hello.c.

  8. Exercise the timeout path. You do not need a hostile file to see it: set a timeout shorter than a worker takes to start.

    bash
    .venv/bin/python triage.py inbox --timeout 0.01 --out reports_t
    text
    intake: 6 files, 5 unique, 1 duplicate(s)
    sha256       type    score  verdict    names                              notes
    78afa14aa60b ?           -  error      copy_of_hello.exe,hello.exe        timeout after 0.01s
    988c0321ddbe ?           -  error      hello_upx.exe                      timeout after 0.01s
    ac62887ad766 ?           -  error      lab_agent.exe                      timeout after 0.01s
    7616f0c28322 ?           -  error      notes.txt                          timeout after 0.01s
    0d5f29750958 ?           -  error      truncated.exe                      timeout after 0.01s

    Every worker was killed and every sample still got a report saying so. The batch finished.

  9. Use the fuzzy hashes. Compare hello.exe with lab_agent.exe and with the packed copy:

    bash
    .venv/bin/python - <<'EOF'
    import json, glob, ppdeep, tlsh
    r = {json.load(open(f))["names"][-1]: json.load(open(f)) for f in glob.glob("reports/*.json")}
    a, b, c = r["hello.exe"], r["lab_agent.exe"], r["hello_upx.exe"]
    print("ssdeep hello/lab_agent", ppdeep.compare(a["ssdeep"], b["ssdeep"]),
          "hello/upx", ppdeep.compare(a["ssdeep"], c["ssdeep"]))
    print("tlsh   hello/lab_agent", tlsh.diff(a["tlsh"], b["tlsh"]),
          "hello/upx", tlsh.diff(a["tlsh"], c["tlsh"]))
    EOF
    text
    ssdeep hello/lab_agent 0 hello/upx 0
    tlsh   hello/lab_agent 36 hello/upx 332

    ssdeep (a similarity score from 0 to 100) finds no resemblance at all. TLSH (a distance, where lower is closer) places the two small mingw-w64 programs close together, because they share most of their runtime code, and the packed copy far away, because packing changes every byte. Neither hash sees through packing. That is why packed samples go to unpacking or detonation before similarity search.

Questions to answer: Which scoring change would stop the benign packed hello_upx.exe from being "suspicious" without hiding a genuinely packed malicious sample, and how would you test that change? What would happen to the batch if analyse ran in the parent process and pefile entered an infinite loop on one file? Why does the pipeline record rules_sha256, and what question does it let you answer six months from now? How would you extend intake to handle a password-protected ZIP without ever writing its members to a directory users can reach? Which fields of these reports would you index first, and which pivot query would you run on lab_agent.exe?

Key takeaways

  • A pipeline automates lookups and bookkeeping (hashing, routing, enrichment, rule scanning, reporting) and hands analysts a ranked queue; verdicts on new samples, rule promotion and external sharing stay with people.
  • Deduplicate by SHA-256 at intake, route on magic bytes rather than extensions, and keep every analyser's output, and its errors, separate.
  • Score additively with a stated reason for every point, let tested family rules decide on their own, and never let a failed stage read as a clean one.
  • Never execute samples on the pipeline host; isolate each analyser, bound its time, memory and output, and treat archives and extracted strings as hostile.
  • Store one JSON record per sample per run with provenance (pipeline, rule and tool versions), and derive STIX, MISP and search-index views from it.
  • Reuse existing platforms (CAPE, Assemblyline, Karton and MWDB, MISP) where they fit, look up hashes rather than uploading files, and measure false positives and coverage so analyst feedback turns into better rules.