Leçon 10.3 · Au-delà de l'EXE· 45 min
PowerShell, JavaScript and VBScript Malware
Deobfuscate malicious PowerShell, JavaScript and VBScript without running them: neutralise eval sinks, decode layers, read script logs, extract IOCs.
Cette leçon n’est disponible qu’en anglais pour le moment.
Objectifs
- Name the Windows hosts that run script malware (powershell.exe, wscript.exe, cscript.exe, mshta.exe) and where such scripts arrive from
- Recognise the common obfuscation families: concatenation, character codes, Base64 and -EncodedCommand, format strings, escape noise and layered eval
- Peel an obfuscated script layer by layer by replacing its execution sink with a print, without running the payload
- Explain what script block logging (event 4104), module logging, AMSI and Sysmon command lines give a defender
- Extract URLs, paths and next-stage hashes from a decoded script and report them
Not every sample is a PE file. A large share of first-stage malware is a few
kilobytes of text: a .js file inside a ZIP attachment, a .vbs dropped by a
shortcut, an .hta fetched from a link, or a single powershell.exe command
line that an EDR captured seconds before the real payload arrived. There is no
header to parse and nothing to disassemble. What stands between you and the
answer is obfuscation, often several layers of it.
Scripts are also the easiest samples to run by accident. Double-clicking a
.js file on Windows does not open an editor; it executes it. This lesson
takes the view of an analyst handed such a script who must work out what it
does, which infrastructure it contacts and what it drops, without letting it
run on a real machine.
Where script malware shows up
Every Windows machine ships with script interpreters, scripts are trivial to change per campaign, and they can fetch and run the next stage without a compiled binary. Typical entry points:
- Email attachments:
.js,.vbs,.wsfor.htafiles, usually inside a ZIP, ISO or password-protected archive to get past mail filters. - Shortcut files (
.lnk) whose target ispowershell.exeormshta.exewith a long argument string. - Office macros that do little more than launch PowerShell with an encoded command. The upcoming Malicious Documents and Archives lesson covers the document side.
- Fake update and CAPTCHA pages that ask the victim to paste a command into the Run dialog or a terminal.
- Post-exploitation: attackers already on a network use PowerShell for discovery, lateral movement and loading tools straight into memory.
In most of these cases the script is a downloader: it fetches a second stage (an EXE, a DLL, shellcode or another script) and runs it. Your job is to find what it fetches, from where, and how it runs it.
The hosts that run them
| Host | Runs | Notes |
|---|---|---|
powershell.exe / pwsh.exe | .ps1, inline -Command, -EncodedCommand | Full access to .NET; Windows PowerShell 5.1 is built in, PowerShell 7 is pwsh.exe |
wscript.exe | .js, .jse, .vbs, .vbe, .wsf | Windows Script Host, GUI flavour; the default handler when a user double-clicks |
cscript.exe | Same as wscript.exe | Console flavour of Windows Script Host; output goes to the terminal |
mshta.exe | .hta, inline vbscript: or javascript: URLs | Runs an HTML application with full local privileges, outside the browser sandbox |
cmd.exe | .bat, .cmd | Mostly used as glue to start one of the hosts above |
Two things surprise newcomers. First, a .js file under Windows Script Host
is JScript, Microsoft's old JavaScript dialect, not browser or Node.js
JavaScript: there is no document and no console, but there is
WScript.Shell for running commands and ActiveXObject for HTTP downloads and
file access. Second, .jse and .vbe are "encoded" scripts, a trivial
Microsoft scheme that public decoders reverse instantly; they are not
encrypted. VBScript has been announced as deprecated by Microsoft, but it is
still present on most systems you will investigate.
The parent process matters as much as the file: winword.exe spawning
powershell.exe is a detection in itself. Processes, Threads and
DLLs explains the process tree.
Obfuscation families
Script obfuscation has one goal: make sure no scanner and no human sees the
interesting strings (http://, DownloadString, Invoke-Expression,
WScript.Shell) in the file as delivered. The techniques are few and they
stack.
| Family | PowerShell example | JavaScript / VBScript example |
|---|---|---|
| Concatenation and splitting | 'Down' + 'loadStr' + 'ing' | "WScr" + "ipt.Sh" + "ell" |
| Character codes | [char]73 + [char]69 + [char]88 | String.fromCharCode(101,118,97,108), Chr(69) & Chr(120) |
| Base64 | [Convert]::FromBase64String(...), -EncodedCommand | atob(...) in browsers; custom decoders in JScript |
| Format-string reordering | "{2}{0}{1}" -f 'ad','String','Downlo' | array lookups by shuffled index: _0x3a[4] + _0x3a[1] |
| Escape and case noise | I`nv`oKe-ExPrEsSiOn (backticks are ignored) | c^m^d in cmd.exe (carets are ignored); random case in VBScript |
| Renaming | $qZx, $____, ${ } | _0x4f2a, single letters, look-alike Unicode |
| Reversal and slicing | -join ('gnirts'[-1..-6]) | .split("").reverse().join("") |
| Encryption or compression | IO.Compression.DeflateStream, XOR loops | XOR or RC4 loops with a key stored beside the blob |
| Layered execution | Invoke-Expression, iex, [scriptblock]::Create() | eval, new Function, VBScript Execute, ExecuteGlobal |
Most of these only change how a string is spelled. The last row is what turns the string back into code, and it is the key to the whole process.
-EncodedCommand is UTF-16LE
PowerShell's -EncodedCommand parameter (accepted as -e, -en, -enc,
-ec and other prefixes, in any case) takes Base64 of the command text
encoded as UTF-16LE, the same wide encoding described in Strings and
Obfuscated Strings. Decode it as UTF-8 and every
character is followed by a zero byte. You can recognise these blobs by eye:
Base64 of UTF-16LE ASCII text is full of A characters, because each zero byte
contributes to runs like AA and patterns such as JABhACAA, the encoding of
$a . Decode with UTF-16LE (CyberChef's From Base64 followed by Decode
text with UTF-16LE) and you get readable code.
Why layers?
Each layer is a string that the previous layer decodes and hands to an
execution sink. A typical chain: a .lnk runs powershell -enc <Base64>; the
decoded command concatenates fragments into a second Base64 blob, decompresses
it with DeflateStream and passes the result to Invoke-Expression; that
third layer downloads a DLL and loads it in memory with
[Reflection.Assembly]::Load, the scripted cousin of reflective DLL
injection. Three layers, and only the
last one does anything interesting.
Peeling a script safely
The method that works on almost everything:
- Work on a copy, in the lab. Rename the file with a harmless extension
(
.js.txt,.ps1.txt) so a stray double-click opens an editor, and do the work in your analysis VM with networking off or simulated. Hash the original first. - Beautify. One-line scripts become readable after
js-beautify, a PowerShell formatter or simply splitting on;. Rename variables as you understand them. - Find the sinks. Search for
eval,Function(,Execute,ExecuteGlobal,Invoke-Expression,iex,ScriptBlock,.Invoke(,Start-Process,WScript.Shell,.Run(,.Exec(andActiveXObject. Everything before a sink is usually decoding; the sink is where the next layer comes to life. - Replace the sink with a print.
eval(x)becomesconsole.log(x),Invoke-Expression $xbecomesWrite-Output $x, VBScriptExecute xbecomesWScript.Echo x. The decoding code runs, the decoded result is printed instead of executed. - Repeat on the printed output until you reach code that does something other than decode code.
- Decode statically when you can. If a layer is just Base64 or a character-code array, decode it in CyberChef or a few lines of Python instead of running anything at all.
Warning: Replacing the final sink is only safe if nothing before it acts on the system. Read every line that will run. Malware often downloads, writes a file or checks the environment early and decodes later. If a line calls
DownloadString,Run,Start-Process,WScript.Shellor anActiveXObject, stub it out too, or decode that part by hand. And run the neutralised script in the lab VM, never on your workstation.
For JScript there is an extra layer of safety: Node.js does not have
WScript or ActiveXObject, so Windows-specific calls fail instead of
running. Emulators such as box-js go further and provide fake versions of
these objects that log every call (URLs requested, files written, commands
run) without performing them.
Tip: Keep every layer. Save each decoded stage as its own file (
stage1.ps1,stage2.ps1...) and note how you got from one to the next. Your report needs the chain, and a later layer sometimes reuses a key or a helper from an earlier one.
Static first, controlled execution second
Static peeling is the default: you see every step, nothing talks to the network, and the decoder works on the next sample of the family. It fails when a layer's key is derived at run time (host name, registry value, server response), or when hand-emulating the obfuscation takes too long.
Then execution becomes the better tool, under control: the real host
(powershell.exe, wscript.exe) inside the isolated VM, with a
sandbox or monitoring tools recording what happens,
fake network services answering, and a snapshot to revert to. This is
dynamic analysis proper; the upcoming
Behavioural Monitoring lesson in Module 6 covers the tooling. For PowerShell
the lab has a special advantage, described next: Windows itself will log the
decoded layers for you.
What Windows telemetry records
Script malware has a weakness binaries do not: the interpreter is Microsoft's, and Microsoft instrumented it. Four sources matter.
| Source | Where | What it captures |
|---|---|---|
| Script block logging | Microsoft-Windows-PowerShell/Operational, event 4104 | The text of every script block PowerShell compiles, including blocks created at run time |
| Module logging | Same log, event 4103 | Pipeline execution details: commands invoked and their parameters |
| AMSI | Delivered to the installed antimalware product | Script content handed to the scanner just before it runs, from PowerShell, Windows Script Host, Office VBA and others |
| Process creation | Sysmon event 1, or Security event 4688 with command-line auditing | Full command line, parent process, hashes (Sysmon), user |
Script block logging is the one analysts love. PowerShell logs a script
block when it compiles it, and every Invoke-Expression or
[scriptblock]::Create() compiles a new block. So when layer 1 decodes layer 2
and passes it to Invoke-Expression, layer 2 is logged as its own 4104 event,
in the form the engine received it: after the Base64 decoding, decompression
and concatenation that produced it. Obfuscation that happens inside a block is
still visible in that block's text, but each layer the attacker executes lands
in the log in clear. Large blocks are split across several events that share a
script block ID. Script block logging is enabled by Group Policy ("Turn on
PowerShell Script Block Logging"); even without it, PowerShell 5 and later log
blocks that contain known suspicious keywords at Warning level.
AMSI, the Antimalware Scan Interface, works on the same principle: the
host passes the content it is about to execute to the antivirus through
amsi.dll, so the scanner sees eval's argument rather than the obfuscated
file. This is why so much malicious PowerShell starts with an AMSI bypass, for
example code that patches AmsiScanBuffer in memory. Strings like
AmsiScanBuffer, amsiInitFailed or System.Management.Automation.AmsiUtils
in a script are a strong signal on their own and good YARA
material.
Process creation events give you the launcher: the exact
powershell.exe -NoP -W Hidden -Enc ... line, its parent and the user. That is
often all an incident responder sends you, and the lab below starts from such a
line.
Tip: In the lab VM, enable script block logging before you run a PowerShell sample, then read the 4104 events afterwards with Event Viewer or
Get-WinEvent -LogName 'Microsoft-Windows-PowerShell/Operational'. You get every decoded layer, in order, without writing a single decoder.
Extracting IOCs and reporting
Once the last layer is readable, collect:
- Network indicators: URLs, domains, IP addresses, user agents and URL paths. Scripts often carry several fallback URLs; take all of them.
- Host indicators: files written (
$env:TEMP\x.dll,%APPDATA%\...\update.vbs), registry keys, scheduled task names, and the command line used to run the next stage (rundll32.exe x.dll,Start,regsvr32 /s). - Next-stage hashes: if the lab served or captured the downloaded file, hash it and triage it as a new sample with the steps from Identifying and Hashing Files.
- Keys and constants used by the decoders, which identify the family or toolkit and make good signature material.
In the report, following the triage report
structure, describe the delivery (attachment, shortcut, macro), the
host (which interpreter and parent process), the layer chain (each
layer's encoding and its hash, labelled as derived artefacts), the final
behaviour, and the IOCs. Add detection advice defenders can act on: the
parent-child pair, the command-line pattern, the 4104 content to hunt for.
Defang every URL (hxxp://update[.]example[.]com/lab) so
nobody clicks it from the report.
Lab: two scripts, peeled without running the payload
You will deobfuscate two small, harmless samples whose real behaviour is only
to print the demo string http://update.example.com/lab. The outputs below are
real, from Node.js v25.7.0 and Python 3.14.7 on macOS. PowerShell was not
installed on that machine, so the PowerShell layer is decoded with Python and
the pwsh equivalent is described but not shown.
-
Create
invoice_0923.js, a JavaScript sample using a character-code array, string reversal andeval:js // invoice_0923.js: harmless lab sample var _0x4f = [59,41,117,40,103,111,108,46,101,108,111,115,110,111,99,32,59,34,98,97,108,47,109,111,99,46,101,108,112,109,97,120,101,34,32,43,32,34,46,101,116,97,100,112,117,47,47,58,112,116,116,104,34,32,61,32,117,32,114,97,118]; var _0x9a = String.fromCharCode.apply(null, _0x4f); var _0x1c = _0x9a.split("").reverse().join(""); eval(_0x1c); -
Create
cmdline.txt, a PowerShell command line as an EDR would record it:text powershell.exe -NoP -W Hidden -Enc JABhACAAPQAgACcAVwByAGkAdABlAC0ATwB1AHQAJwAgACsAIAAnAHAAdQB0ACcAOwAgACQAYgAgAD0AIAAnACgAJwAnAGgAdAB0AHAAOgAvAC8AdQBwAGQAJwAnACAAKwAgACcAJwBhAHQAZQAuAGUAeABhAG0AcABsAGUALgBjAG8AbQAvAGwAYQBiACcAJwApACcAOwAgAEkAbgB2AG8AawBlAC0ARQB4AHAAcgBlAHMAcwBpAG8AbgAgACgAJABhACAAKwAgACcAIAAnACAAKwAgACQAYgApAA== -
Confirm that the indicator is not visible, then find the sinks:
bash grep -c example invoice_0923.js cmdline.txt grep -nE 'eval|Function\(|Invoke-Expression|IEX|-Enc' invoice_0923.js cmdline.txt | cut -c1-90text invoice_0923.js:0 cmdline.txt:0 invoice_0923.js:5:eval(_0x1c); cmdline.txt:1:powershell.exe -NoP -W Hidden -Enc JABhACAAPQAgACcAVwByAGkAdABlAC0ATwB1AHQAJNeither file contains the domain. The JavaScript sink is
evalon line 5. The PowerShell sample has no visible sink yet:-Enchides the whole command. -
Read the JavaScript before running anything. Lines 2 to 4 only build a string:
fromCharCodeturns numbers into characters, thensplit,reverseandjoinreverse it. Nothing touches the file system or the network untileval. Neutralise the sink and check the change:bash sed 's/^eval(/console.log(/' invoice_0923.js > invoice_0923.safe.js diff invoice_0923.js invoice_0923.safe.jstext 5c5 < eval(_0x1c); --- > console.log(_0x1c); -
Run only the neutralised copy (in your lab VM):
bash node invoice_0923.safe.jstext var u = "http://update." + "example.com/lab"; console.log(u);The printed text is layer 2, as source code: a concatenation that builds the URL and prints it. A real sample would call
new ActiveXObject("MSXML2.XMLHTTP")here instead ofconsole.log. Layer 2 has no further sink, so peeling stops; the concatenation is simple enough to join by eye. You now have the URL without executing it. -
Try decoding the PowerShell blob the naive way:
bash awk '{print $NF}' cmdline.txt | base64 -d | head -c 60 | xxd | head -3text 00000000: 2400 6100 2000 3d00 2000 2700 5700 7200 $.a. .=. .'.W.r. 00000010: 6900 7400 6500 2d00 4f00 7500 7400 2700 i.t.e.-.O.u.t.'. 00000020: 2000 2b00 2000 2700 7000 7500 7400 2700 .+. .'.p.u.t.'.Every other byte is zero: this is UTF-16LE text, exactly as
-EncodedCommandrequires. -
Save
decode_enc.py, which decodes the argument as UTF-16LE, folds'a' + 'b'concatenations of single-quoted literals, and then treats each literal as potential code for the next layer:python import base64, re, sys # Tokens: a whole single-quoted literal ('' inside is an escaped quote), # a plus sign, or a run of anything else. TOKEN = re.compile(r"'(?:[^']|'')*'|\+|[^'+]+") def fold(code): """Join 'abc' + 'def' into 'abcdef', scanning left to right.""" out = [] for t in TOKEN.findall(code): t = t.strip() if not t.startswith("'") else t if not t: continue # top of the output stack is <literal> + : merge the two literals if t.startswith("'") and out[-2:-1] and out[-1] == "+" and out[-2].startswith("'"): out.pop() t = out.pop()[:-1] + t[1:] out.append(t) return " ".join(out) def literals(code): """Contents of each literal, with '' turned back into '.""" return [t[1:-1].replace("''", "'") for t in TOKEN.findall(code) if t.startswith("'")] cmdline = open(sys.argv[1], encoding="utf-8").read().strip() # 1. The argument after -e / -enc / -EncodedCommand (PowerShell accepts prefixes). blob = re.search(r"-e[a-z]*\s+([A-Za-z0-9+/=]+)", cmdline, re.I).group(1) print(f"[+] Base64 argument: {len(blob)} chars") # 2. -EncodedCommand is Base64 of UTF-16LE text, not UTF-8. layer1 = base64.b64decode(blob).decode("utf-16le") print("[+] layer 1, decoded:\n " + layer1) print("[+] layer 1, folded:\n " + fold(layer1)) # 3. A literal may itself be code for Invoke-Expression: unescape and fold again. print("[+] layer 2, literals folded:") urls = set() for s in literals(fold(layer1)): inner = fold(s) print(" " + inner) urls.update(re.findall(r"https?://[^\s'\")]+", inner)) print("[+] URLs:", sorted(urls))The tokeniser consumes each literal whole from its opening quote, because
''inside it is an escaped quote. A simpler search-and-replace regex joined fragments inside$band corrupted the URL. Decoders fail quietly: always compare their output with the input by eye. -
Run it:
bash python3 decode_enc.py cmdline.txttext [+] Base64 argument: 296 chars [+] layer 1, decoded: $a = 'Write-Out' + 'put'; $b = '(''http://upd'' + ''ate.example.com/lab'')'; Invoke-Expression ($a + ' ' + $b) [+] layer 1, folded: $a = 'Write-Output' ; $b = '(''http://upd'' + ''ate.example.com/lab'')' ; Invoke-Expression ($a + ' ' + $b) [+] layer 2, literals folded: Write-Output ( 'http://update.example.com/lab' ) [+] URLs: ['http://update.example.com/lab']Layer 1 builds a command name and an argument from fragments and passes them to
Invoke-Expression. The string$bholds layer 2 as source code, with its own concatenation; folding it once more reveals the URL. The empty line is the' 'separator literal. Nothing was executed at any point. -
With PowerShell available (inside the lab VM,
pwshon any platform or Windows PowerShell), the same peeling uses the interpreter itself for the harmless parts. Decode the argument with[Text.Encoding]::Unicode.GetString([Convert]::FromBase64String($blob))(Unicodeis .NET's name for UTF-16LE) and save it aslayer1.ps1. Read it, confirm that only string operations precede the sink, replaceInvoke-ExpressionwithWrite-Outputand run the edited copy withpwsh -NoProfile -File layer1.ps1. It should print layer 2,Write-Output ('http://upd' + 'ate.example.com/lab'), as text. If you instead run the original command line in a Windows lab VM with script block logging enabled, event 4104 records both layer 1 and the block created byInvoke-Expression. -
Paste the Base64 argument into CyberChef, apply From Base64 then Decode text (UTF-16LE (1200)), and compare with step 8.
Questions to answer: Why does grep find neither the domain nor
Invoke-Expression in cmdline.txt? In step 4, which lines would you have to
stub out if line 3 had been new ActiveXObject("WScript.Shell").Run(...)
instead of a string operation? If the sample had used
"{1}{0}" -f 'ssion','Invoke-Expre' instead of 'Invoke-Expre' + 'ssion', how
would decode_enc.py need to change? Which 4104 events would a defender see if
this command line ran on a machine with script block logging enabled, and what
would each contain? Write the IOC section of a report for these two samples,
with the URL defanged.
Key takeaways
- Script malware runs on hosts every Windows machine ships with:
powershell.exe,wscript.exeandcscript.exefor JScript and VBScript, andmshta.exefor HTML applications. The parent process is part of the evidence. - Obfuscation changes how strings are spelled (concatenation, character codes,
Base64, format strings, escape noise, renaming); an execution sink such as
evalorInvoke-Expressionturns the result into the next layer. -EncodedCommandis Base64 of UTF-16LE. Decode it as UTF-16LE, in CyberChef or three lines of Python.- Peel by replacing the sink with a print, after checking that nothing before it acts on the system, and repeat until the code stops decoding code. Decode statically whenever you can, and run anything only in the lab.
- Script block logging (event 4104) records each script block as PowerShell
compiles it, so every layer passed to
Invoke-Expressionappears in clear; AMSI gives antimalware the same view, which is why attackers try to disable it. - Report the delivery, host, layer chain with hashes, final behaviour and defanged IOCs, plus the command-line and parent-process patterns defenders can hunt for.