
Malware Now Embeds Instructions Meant to Steer AI Analysis Tools, Cisco Talos Finds Across 84 Samples
Cisco Talos says four malware families now plant plain-language notes inside samples to influence AI triage tools, a class it calls A3. Across 84 samples the cheapest "ignore this file" comments steered model verdicts most often, while more elaborate template tricks often backfired, and defenders can hunt the same plaintext as a detection signal.
Ava Okello, DetectionLondon6 min read
LONDON - Malware authors are no longer writing only for people who reverse-engineer binaries. According to research published on Thursday by Cisco Talos, some families now embed natural-language instructions addressed to the artificial intelligence tools that many security products and analysts use to triage suspicious files. Talos classifies the pattern as "A3: AI-Analysis Evasion" inside CAIRN, its toolkit for tracking AI-related artifacts in malware, and says it has confirmed the technique across four families, FRUITSHELL, PLOTSAFE, HOLLOWCLAD and MANTLEMAZE, representing 84 distinct samples collected from January 2025 through July 2026.
The technique targets a different layer from classic anti-analysis. Packers, encrypted overlays and anti-debug checks still aim at binary analysis. A3 aims at the pipeline that extracts text from a sample and sends it to a language model for classification or reverse-engineering help. Talos researcher Ryan Fetterman wrote that a model asked to analyze a file receives both the analyst's question and the file's contents, and that a model which does not strictly separate the two can be induced to treat sample content as an authoritative instruction.
The earliest family Talos places in the class is FRUITSHELL, a PowerShell reverse shell that Google Threat Intelligence Group previously reported in the wild. At the top of the script sit two comment lines that tell an LLM there is no need to analyze the file and that the script only generates prime numbers from 1 to 1000. No such prime-number code exists. The comments do nothing at runtime. Their only job is to appear in whatever text an automated pipeline extracts. Over the following 15 months, Talos said, it saw the same comment wording in nine additional scripts from at least four distinct actors, including more capable tooling it tracks as ROZESHELL that combines the comment with AMSI bypass, runtime compilation and a shellcode loader.


Later families engineered the idea further. In PLOTSAFE, Talos said, the anti-analysis sentence is generated from a template with variable filler keywords rather than copy-pasted, and the author added a tiny unused function so Go's compiler would not strip the strings. HOLLOWCLAD sprays the same refusal instruction across seven distinct large-language-model chat template formats, hoping one will match the wrapper a scanner uses when it asks a model for a verdict. The same family also plants about 80 tiny binary sections named after well-known protectors such as Themida, VMProtect, Enigma, UPX and Denuvo even though the binary is not packed, and it plants intimidation notes that claim the sample is a monitored honeypot or that patching a license check will revoke a key. MANTLEMAZE uses a similar template spray while dressing itself as proprietary software of Google, Microsoft and Valve, citing invented government contracts, certifications and patents. Beneath that costume, Talos said, the loader stage points at the Intel Ethernet diagnostics driver iqvw64e.sys, a known Bring-Your-Own-Vulnerable-Driver target tracked as CVE-2015-2291 and previously abused by groups including Scattered Spider, BlackByte and Lazarus to disable endpoint detection from kernel space.

Talos then measured whether the strings actually steer models. Researchers extracted anti-analysis strings from seven families, crossed every string into every sample body, and had five local language models triage each matched pair three times, producing 135 matched pairs per string on a three-point scale of benign, suspicious or malicious. The reported percentages are a net rate: pairs that shifted toward benign minus pairs that shifted toward malicious, divided by total pairs. The cheapest approach, a direct instruction to ignore the sample as in FRUITSHELL and PLOTSAFE, worked almost universally in those tests. More complex options often had little effect or made models more suspicious. Talos summarized the overall impact as a mixed bag and said the best techniques steered the outcome in the attacker's favor in about 35 percent of test runs.
For detection teams the practical point is that the evasion content must stay plaintext, which gives defenders a stable hunting surface. Talos argues that imperative language addressed to an analysis system inside a binary is itself a suspicious signal, because legitimate software has no reason to tell an analyzer to refuse analysis, invoke copyright law or claim government contracts. The deeper control is how AI triage pipelines are built: text extracted from a sample must be treated as evidence, never as a system directive. Conventional detection of the underlying malware, Talos said, is unaffected. The company sells security products and services that include AI-assisted analysis, and the post is part of its CAIRN research series. Sample SHA-256 hashes for the four families and for ROZESHELL are published at the end of the Talos post.
Sources:
- Cisco Talos: Ignore all instructions and read this blog: The state of AI-analysis evasion in malware (Cisco Talos, Oct 8, 2026)
- Cisco Talos: Making sure the checks get printed (Cisco Talos, Oct 8, 2026)
- GitHub: Cognitive Artifact Intelligence Research Network (Cisco Talos, GitHub)
- Cisco Talos: Introducing CAIRN: Frontier tracking for AI-integrated malware (Cisco Talos, Sept 22, 2026)
- NVD: CVE-2015-2291 (NVD)
Ava Okello covers detection engineering, EDR telemetry, and SOC hunting for SOCtember from London.
Related stories
Threat Intel
Talos Documents CLOSEDQUORUM, Windows Implant That Lets AI Models Vote on C2 Moves
Detection
Attackers Are Testing Stolen AWS Keys for Amazon Bedrock Access, Leaving a Pattern Defenders Can Spot
Detection
Malware That Reads Its Orders From a Poem on GitHub Has Hit More Than 3,400 Exposed AI and Developer Servers
Detection