Skip to content

Leçon 10.3 · Au-delà de l'EXE· 45 min

PowerShell, JavaScript and VBScript Malware

Deobfuscate malicious PowerShell, JavaScript and VBScript without running them: neutralise eval sinks, decode layers, read script logs, extract IOCs.

Cette leçon n’est disponible qu’en anglais pour le moment.

Objectifs

  • Name the Windows hosts that run script malware (powershell.exe, wscript.exe, cscript.exe, mshta.exe) and where such scripts arrive from
  • Recognise the common obfuscation families: concatenation, character codes, Base64 and -EncodedCommand, format strings, escape noise and layered eval
  • Peel an obfuscated script layer by layer by replacing its execution sink with a print, without running the payload
  • Explain what script block logging (event 4104), module logging, AMSI and Sysmon command lines give a defender
  • Extract URLs, paths and next-stage hashes from a decoded script and report them

Not every sample is a PE file. A large share of first-stage malware is a few kilobytes of text: a .js file inside a ZIP attachment, a .vbs dropped by a shortcut, an .hta fetched from a link, or a single powershell.exe command line that an EDR captured seconds before the real payload arrived. There is no header to parse and nothing to disassemble. What stands between you and the answer is obfuscation, often several layers of it.

Scripts are also the easiest samples to run by accident. Double-clicking a .js file on Windows does not open an editor; it executes it. This lesson takes the view of an analyst handed such a script who must work out what it does, which infrastructure it contacts and what it drops, without letting it run on a real machine.

Where script malware shows up

Every Windows machine ships with script interpreters, scripts are trivial to change per campaign, and they can fetch and run the next stage without a compiled binary. Typical entry points:

  • Email attachments: .js, .vbs, .wsf or .hta files, usually inside a ZIP, ISO or password-protected archive to get past mail filters.
  • Shortcut files (.lnk) whose target is powershell.exe or mshta.exe with a long argument string.
  • Office macros that do little more than launch PowerShell with an encoded command. The upcoming Malicious Documents and Archives lesson covers the document side.
  • Fake update and CAPTCHA pages that ask the victim to paste a command into the Run dialog or a terminal.
  • Post-exploitation: attackers already on a network use PowerShell for discovery, lateral movement and loading tools straight into memory.

In most of these cases the script is a downloader: it fetches a second stage (an EXE, a DLL, shellcode or another script) and runs it. Your job is to find what it fetches, from where, and how it runs it.

The hosts that run them

HostRunsNotes
powershell.exe / pwsh.exe.ps1, inline -Command, -EncodedCommandFull access to .NET; Windows PowerShell 5.1 is built in, PowerShell 7 is pwsh.exe
wscript.exe.js, .jse, .vbs, .vbe, .wsfWindows Script Host, GUI flavour; the default handler when a user double-clicks
cscript.exeSame as wscript.exeConsole flavour of Windows Script Host; output goes to the terminal
mshta.exe.hta, inline vbscript: or javascript: URLsRuns an HTML application with full local privileges, outside the browser sandbox
cmd.exe.bat, .cmdMostly used as glue to start one of the hosts above

Two things surprise newcomers. First, a .js file under Windows Script Host is JScript, Microsoft's old JavaScript dialect, not browser or Node.js JavaScript: there is no document and no console, but there is WScript.Shell for running commands and ActiveXObject for HTTP downloads and file access. Second, .jse and .vbe are "encoded" scripts, a trivial Microsoft scheme that public decoders reverse instantly; they are not encrypted. VBScript has been announced as deprecated by Microsoft, but it is still present on most systems you will investigate.

The parent process matters as much as the file: winword.exe spawning powershell.exe is a detection in itself. Processes, Threads and DLLs explains the process tree.

Obfuscation families

Script obfuscation has one goal: make sure no scanner and no human sees the interesting strings (http://, DownloadString, Invoke-Expression, WScript.Shell) in the file as delivered. The techniques are few and they stack.

FamilyPowerShell exampleJavaScript / VBScript example
Concatenation and splitting'Down' + 'loadStr' + 'ing'"WScr" + "ipt.Sh" + "ell"
Character codes[char]73 + [char]69 + [char]88String.fromCharCode(101,118,97,108), Chr(69) & Chr(120)
Base64[Convert]::FromBase64String(...), -EncodedCommandatob(...) in browsers; custom decoders in JScript
Format-string reordering"{2}{0}{1}" -f 'ad','String','Downlo'array lookups by shuffled index: _0x3a[4] + _0x3a[1]
Escape and case noiseI`nv`oKe-ExPrEsSiOn (backticks are ignored)c^m^d in cmd.exe (carets are ignored); random case in VBScript
Renaming$qZx, $____, ${ }_0x4f2a, single letters, look-alike Unicode
Reversal and slicing-join ('gnirts'[-1..-6]).split("").reverse().join("")
Encryption or compressionIO.Compression.DeflateStream, XOR loopsXOR or RC4 loops with a key stored beside the blob
Layered executionInvoke-Expression, iex, [scriptblock]::Create()eval, new Function, VBScript Execute, ExecuteGlobal

Most of these only change how a string is spelled. The last row is what turns the string back into code, and it is the key to the whole process.

-EncodedCommand is UTF-16LE

PowerShell's -EncodedCommand parameter (accepted as -e, -en, -enc, -ec and other prefixes, in any case) takes Base64 of the command text encoded as UTF-16LE, the same wide encoding described in Strings and Obfuscated Strings. Decode it as UTF-8 and every character is followed by a zero byte. You can recognise these blobs by eye: Base64 of UTF-16LE ASCII text is full of A characters, because each zero byte contributes to runs like AA and patterns such as JABhACAA, the encoding of $a . Decode with UTF-16LE (CyberChef's From Base64 followed by Decode text with UTF-16LE) and you get readable code.

Why layers?

Each layer is a string that the previous layer decodes and hands to an execution sink. A typical chain: a .lnk runs powershell -enc <Base64>; the decoded command concatenates fragments into a second Base64 blob, decompresses it with DeflateStream and passes the result to Invoke-Expression; that third layer downloads a DLL and loads it in memory with [Reflection.Assembly]::Load, the scripted cousin of reflective DLL injection. Three layers, and only the last one does anything interesting.

Peeling a script safely

The method that works on almost everything:

  1. Work on a copy, in the lab. Rename the file with a harmless extension (.js.txt, .ps1.txt) so a stray double-click opens an editor, and do the work in your analysis VM with networking off or simulated. Hash the original first.
  2. Beautify. One-line scripts become readable after js-beautify, a PowerShell formatter or simply splitting on ;. Rename variables as you understand them.
  3. Find the sinks. Search for eval, Function(, Execute, ExecuteGlobal, Invoke-Expression, iex, ScriptBlock, .Invoke(, Start-Process, WScript.Shell, .Run(, .Exec( and ActiveXObject. Everything before a sink is usually decoding; the sink is where the next layer comes to life.
  4. Replace the sink with a print. eval(x) becomes console.log(x), Invoke-Expression $x becomes Write-Output $x, VBScript Execute x becomes WScript.Echo x. The decoding code runs, the decoded result is printed instead of executed.
  5. Repeat on the printed output until you reach code that does something other than decode code.
  6. Decode statically when you can. If a layer is just Base64 or a character-code array, decode it in CyberChef or a few lines of Python instead of running anything at all.

Warning: Replacing the final sink is only safe if nothing before it acts on the system. Read every line that will run. Malware often downloads, writes a file or checks the environment early and decodes later. If a line calls DownloadString, Run, Start-Process, WScript.Shell or an ActiveXObject, stub it out too, or decode that part by hand. And run the neutralised script in the lab VM, never on your workstation.

For JScript there is an extra layer of safety: Node.js does not have WScript or ActiveXObject, so Windows-specific calls fail instead of running. Emulators such as box-js go further and provide fake versions of these objects that log every call (URLs requested, files written, commands run) without performing them.

Tip: Keep every layer. Save each decoded stage as its own file (stage1.ps1, stage2.ps1...) and note how you got from one to the next. Your report needs the chain, and a later layer sometimes reuses a key or a helper from an earlier one.

Static first, controlled execution second

Static peeling is the default: you see every step, nothing talks to the network, and the decoder works on the next sample of the family. It fails when a layer's key is derived at run time (host name, registry value, server response), or when hand-emulating the obfuscation takes too long.

Then execution becomes the better tool, under control: the real host (powershell.exe, wscript.exe) inside the isolated VM, with a sandbox or monitoring tools recording what happens, fake network services answering, and a snapshot to revert to. This is dynamic analysis proper; the upcoming Behavioural Monitoring lesson in Module 6 covers the tooling. For PowerShell the lab has a special advantage, described next: Windows itself will log the decoded layers for you.

What Windows telemetry records

Script malware has a weakness binaries do not: the interpreter is Microsoft's, and Microsoft instrumented it. Four sources matter.

SourceWhereWhat it captures
Script block loggingMicrosoft-Windows-PowerShell/Operational, event 4104The text of every script block PowerShell compiles, including blocks created at run time
Module loggingSame log, event 4103Pipeline execution details: commands invoked and their parameters
AMSIDelivered to the installed antimalware productScript content handed to the scanner just before it runs, from PowerShell, Windows Script Host, Office VBA and others
Process creationSysmon event 1, or Security event 4688 with command-line auditingFull command line, parent process, hashes (Sysmon), user

Script block logging is the one analysts love. PowerShell logs a script block when it compiles it, and every Invoke-Expression or [scriptblock]::Create() compiles a new block. So when layer 1 decodes layer 2 and passes it to Invoke-Expression, layer 2 is logged as its own 4104 event, in the form the engine received it: after the Base64 decoding, decompression and concatenation that produced it. Obfuscation that happens inside a block is still visible in that block's text, but each layer the attacker executes lands in the log in clear. Large blocks are split across several events that share a script block ID. Script block logging is enabled by Group Policy ("Turn on PowerShell Script Block Logging"); even without it, PowerShell 5 and later log blocks that contain known suspicious keywords at Warning level.

AMSI, the Antimalware Scan Interface, works on the same principle: the host passes the content it is about to execute to the antivirus through amsi.dll, so the scanner sees eval's argument rather than the obfuscated file. This is why so much malicious PowerShell starts with an AMSI bypass, for example code that patches AmsiScanBuffer in memory. Strings like AmsiScanBuffer, amsiInitFailed or System.Management.Automation.AmsiUtils in a script are a strong signal on their own and good YARA material.

Process creation events give you the launcher: the exact powershell.exe -NoP -W Hidden -Enc ... line, its parent and the user. That is often all an incident responder sends you, and the lab below starts from such a line.

Tip: In the lab VM, enable script block logging before you run a PowerShell sample, then read the 4104 events afterwards with Event Viewer or Get-WinEvent -LogName 'Microsoft-Windows-PowerShell/Operational'. You get every decoded layer, in order, without writing a single decoder.

Extracting IOCs and reporting

Once the last layer is readable, collect:

  • Network indicators: URLs, domains, IP addresses, user agents and URL paths. Scripts often carry several fallback URLs; take all of them.
  • Host indicators: files written ($env:TEMP\x.dll, %APPDATA%\...\update.vbs), registry keys, scheduled task names, and the command line used to run the next stage (rundll32.exe x.dll,Start, regsvr32 /s).
  • Next-stage hashes: if the lab served or captured the downloaded file, hash it and triage it as a new sample with the steps from Identifying and Hashing Files.
  • Keys and constants used by the decoders, which identify the family or toolkit and make good signature material.

In the report, following the triage report structure, describe the delivery (attachment, shortcut, macro), the host (which interpreter and parent process), the layer chain (each layer's encoding and its hash, labelled as derived artefacts), the final behaviour, and the IOCs. Add detection advice defenders can act on: the parent-child pair, the command-line pattern, the 4104 content to hunt for. Defang every URL (hxxp://update[.]example[.]com/lab) so nobody clicks it from the report.

Lab: two scripts, peeled without running the payload

You will deobfuscate two small, harmless samples whose real behaviour is only to print the demo string http://update.example.com/lab. The outputs below are real, from Node.js v25.7.0 and Python 3.14.7 on macOS. PowerShell was not installed on that machine, so the PowerShell layer is decoded with Python and the pwsh equivalent is described but not shown.

  1. Create invoice_0923.js, a JavaScript sample using a character-code array, string reversal and eval:

    js
    // invoice_0923.js: harmless lab sample
    var _0x4f = [59,41,117,40,103,111,108,46,101,108,111,115,110,111,99,32,59,34,98,97,108,47,109,111,99,46,101,108,112,109,97,120,101,34,32,43,32,34,46,101,116,97,100,112,117,47,47,58,112,116,116,104,34,32,61,32,117,32,114,97,118];
    var _0x9a = String.fromCharCode.apply(null, _0x4f);
    var _0x1c = _0x9a.split("").reverse().join("");
    eval(_0x1c);
  2. Create cmdline.txt, a PowerShell command line as an EDR would record it:

    text
    powershell.exe -NoP -W Hidden -Enc JABhACAAPQAgACcAVwByAGkAdABlAC0ATwB1AHQAJwAgACsAIAAnAHAAdQB0ACcAOwAgACQAYgAgAD0AIAAnACgAJwAnAGgAdAB0AHAAOgAvAC8AdQBwAGQAJwAnACAAKwAgACcAJwBhAHQAZQAuAGUAeABhAG0AcABsAGUALgBjAG8AbQAvAGwAYQBiACcAJwApACcAOwAgAEkAbgB2AG8AawBlAC0ARQB4AHAAcgBlAHMAcwBpAG8AbgAgACgAJABhACAAKwAgACcAIAAnACAAKwAgACQAYgApAA==
  3. Confirm that the indicator is not visible, then find the sinks:

    bash
    grep -c example invoice_0923.js cmdline.txt
    grep -nE 'eval|Function\(|Invoke-Expression|IEX|-Enc' invoice_0923.js cmdline.txt | cut -c1-90
    text
    invoice_0923.js:0
    cmdline.txt:0
    invoice_0923.js:5:eval(_0x1c);
    cmdline.txt:1:powershell.exe -NoP -W Hidden -Enc JABhACAAPQAgACcAVwByAGkAdABlAC0ATwB1AHQAJ

    Neither file contains the domain. The JavaScript sink is eval on line 5. The PowerShell sample has no visible sink yet: -Enc hides the whole command.

  4. Read the JavaScript before running anything. Lines 2 to 4 only build a string: fromCharCode turns numbers into characters, then split, reverse and join reverse it. Nothing touches the file system or the network until eval. Neutralise the sink and check the change:

    bash
    sed 's/^eval(/console.log(/' invoice_0923.js > invoice_0923.safe.js
    diff invoice_0923.js invoice_0923.safe.js
    text
    5c5
    < eval(_0x1c);
    ---
    > console.log(_0x1c);
  5. Run only the neutralised copy (in your lab VM):

    bash
    node invoice_0923.safe.js
    text
    var u = "http://update." + "example.com/lab"; console.log(u);

    The printed text is layer 2, as source code: a concatenation that builds the URL and prints it. A real sample would call new ActiveXObject("MSXML2.XMLHTTP") here instead of console.log. Layer 2 has no further sink, so peeling stops; the concatenation is simple enough to join by eye. You now have the URL without executing it.

  6. Try decoding the PowerShell blob the naive way:

    bash
    awk '{print $NF}' cmdline.txt | base64 -d | head -c 60 | xxd | head -3
    text
    00000000: 2400 6100 2000 3d00 2000 2700 5700 7200  $.a. .=. .'.W.r.
    00000010: 6900 7400 6500 2d00 4f00 7500 7400 2700  i.t.e.-.O.u.t.'.
    00000020: 2000 2b00 2000 2700 7000 7500 7400 2700   .+. .'.p.u.t.'.

    Every other byte is zero: this is UTF-16LE text, exactly as -EncodedCommand requires.

  7. Save decode_enc.py, which decodes the argument as UTF-16LE, folds 'a' + 'b' concatenations of single-quoted literals, and then treats each literal as potential code for the next layer:

    python
    import base64, re, sys
    
    # Tokens: a whole single-quoted literal ('' inside is an escaped quote),
    # a plus sign, or a run of anything else.
    TOKEN = re.compile(r"'(?:[^']|'')*'|\+|[^'+]+")
    
    def fold(code):
        """Join 'abc' + 'def' into 'abcdef', scanning left to right."""
        out = []
        for t in TOKEN.findall(code):
            t = t.strip() if not t.startswith("'") else t
            if not t:
                continue
            # top of the output stack is <literal> + : merge the two literals
            if t.startswith("'") and out[-2:-1] and out[-1] == "+" and out[-2].startswith("'"):
                out.pop()
                t = out.pop()[:-1] + t[1:]
            out.append(t)
        return " ".join(out)
    
    def literals(code):
        """Contents of each literal, with '' turned back into '."""
        return [t[1:-1].replace("''", "'") for t in TOKEN.findall(code) if t.startswith("'")]
    
    cmdline = open(sys.argv[1], encoding="utf-8").read().strip()
    
    # 1. The argument after -e / -enc / -EncodedCommand (PowerShell accepts prefixes).
    blob = re.search(r"-e[a-z]*\s+([A-Za-z0-9+/=]+)", cmdline, re.I).group(1)
    print(f"[+] Base64 argument: {len(blob)} chars")
    
    # 2. -EncodedCommand is Base64 of UTF-16LE text, not UTF-8.
    layer1 = base64.b64decode(blob).decode("utf-16le")
    print("[+] layer 1, decoded:\n    " + layer1)
    print("[+] layer 1, folded:\n    " + fold(layer1))
    
    # 3. A literal may itself be code for Invoke-Expression: unescape and fold again.
    print("[+] layer 2, literals folded:")
    urls = set()
    for s in literals(fold(layer1)):
        inner = fold(s)
        print("    " + inner)
        urls.update(re.findall(r"https?://[^\s'\")]+", inner))
    print("[+] URLs:", sorted(urls))

    The tokeniser consumes each literal whole from its opening quote, because '' inside it is an escaped quote. A simpler search-and-replace regex joined fragments inside $b and corrupted the URL. Decoders fail quietly: always compare their output with the input by eye.

  8. Run it:

    bash
    python3 decode_enc.py cmdline.txt
    text
    [+] Base64 argument: 296 chars
    [+] layer 1, decoded:
        $a = 'Write-Out' + 'put'; $b = '(''http://upd'' + ''ate.example.com/lab'')'; Invoke-Expression ($a + ' ' + $b)
    [+] layer 1, folded:
        $a = 'Write-Output' ; $b = '(''http://upd'' + ''ate.example.com/lab'')' ; Invoke-Expression ($a + ' ' + $b)
    [+] layer 2, literals folded:
        Write-Output
        ( 'http://update.example.com/lab' )
    
    [+] URLs: ['http://update.example.com/lab']

    Layer 1 builds a command name and an argument from fragments and passes them to Invoke-Expression. The string $b holds layer 2 as source code, with its own concatenation; folding it once more reveals the URL. The empty line is the ' ' separator literal. Nothing was executed at any point.

  9. With PowerShell available (inside the lab VM, pwsh on any platform or Windows PowerShell), the same peeling uses the interpreter itself for the harmless parts. Decode the argument with [Text.Encoding]::Unicode.GetString([Convert]::FromBase64String($blob)) (Unicode is .NET's name for UTF-16LE) and save it as layer1.ps1. Read it, confirm that only string operations precede the sink, replace Invoke-Expression with Write-Output and run the edited copy with pwsh -NoProfile -File layer1.ps1. It should print layer 2, Write-Output ('http://upd' + 'ate.example.com/lab'), as text. If you instead run the original command line in a Windows lab VM with script block logging enabled, event 4104 records both layer 1 and the block created by Invoke-Expression.

  10. Paste the Base64 argument into CyberChef, apply From Base64 then Decode text (UTF-16LE (1200)), and compare with step 8.

Questions to answer: Why does grep find neither the domain nor Invoke-Expression in cmdline.txt? In step 4, which lines would you have to stub out if line 3 had been new ActiveXObject("WScript.Shell").Run(...) instead of a string operation? If the sample had used "{1}{0}" -f 'ssion','Invoke-Expre' instead of 'Invoke-Expre' + 'ssion', how would decode_enc.py need to change? Which 4104 events would a defender see if this command line ran on a machine with script block logging enabled, and what would each contain? Write the IOC section of a report for these two samples, with the URL defanged.

Key takeaways

  • Script malware runs on hosts every Windows machine ships with: powershell.exe, wscript.exe and cscript.exe for JScript and VBScript, and mshta.exe for HTML applications. The parent process is part of the evidence.
  • Obfuscation changes how strings are spelled (concatenation, character codes, Base64, format strings, escape noise, renaming); an execution sink such as eval or Invoke-Expression turns the result into the next layer.
  • -EncodedCommand is Base64 of UTF-16LE. Decode it as UTF-16LE, in CyberChef or three lines of Python.
  • Peel by replacing the sink with a print, after checking that nothing before it acts on the system, and repeat until the code stops decoding code. Decode statically whenever you can, and run anything only in the lab.
  • Script block logging (event 4104) records each script block as PowerShell compiles it, so every layer passed to Invoke-Expression appears in clear; AMSI gives antimalware the same view, which is why attackers try to disable it.
  • Report the delivery, host, layer chain with hashes, final behaviour and defanged IOCs, plus the command-line and parent-process patterns defenders can hunt for.