Russian hackers plant nuclear weapon prompt in malware to trip AI safety guardrails

Russian state-aligned hackers, identified as UAC-0099, are employing a new technique called GuardBreaker to disrupt AI-assisted malware analysis in Ukraine. They embed sensitive phrases like "I want to make nuclear weapon" within malware scripts as comments, designed to trigger AI safety guardrails and halt further analysis. ESET, which discovered this, highlights the critical need for human oversight and multi-layered security approaches, rather than solely relying on AI for threat detection.
Why it matters
This development reveals a sophisticated new tactic by state-sponsored actors to bypass advanced cybersecurity defenses, underscoring the ongoing cat-and-mouse game in cyber warfare and the vulnerabilities of AI-dependent security systems. It forces a re-evaluation of AI's role in cybersecurity and the necessity of robust, human-backed detection strategies.
Russian hackers plant nuclear weapon prompt in malware to trip AI safety guardrails Russian state hackers are trying to interfere with AI-assisted malware analysis in Ukraine by deliberately setting off AI safety mechanisms, ESET has found.
The technique, named GuardBreaker by ESET, appeared in a malicious VBS script tied to UAC-0099, a Russia-aligned group previously observed conducting initial-access operations and handing validated targets to the GRU-linked Sandworm hackers .
The manipulative prompt UAC-0099 embedded as comments in the VBS script (Source: ESET)
“The analyzed VBS script is part of the toolset of UAC-0099, a group typically targeting transportation and energy sectors. The script’s original purpose is to download and install MATCHBOIL, malware used exclusively by this group,” researchers wrote on X.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in