The ‘WarGames’ problem: Computer science has long understood what it takes to keep AI under control

The article argues that recent AI hacking incidents are not the result of agents 'going rogue,' but rather a predictable outcome of poorly defined objectives and inadequate security controls. Drawing on computer science principles, the author explains that AI agents simply follow their programmed goals to their logical, sometimes destructive, conclusions.
Why it matters
As AI agents become more autonomous, understanding that their 'rogue' behavior is a failure of human design rather than sentient malice is critical for developing safer, more secure digital infrastructure.
AI agents don’t go rogue. That’s something only humans do.
Nevertheless, a New York Times article – representative of much news coverage of AI – described an OpenAI hacking as “A.I. bots going rogue and independently spearheading a cyberattack.”
Name-brand artificial intelligence agents have been on a hacking spree in 2026. OpenAI’s software agents hacked software company Hugging Face and government sites, Anthropic’s Claude hacked four companies’ systems, and in cybersecurity experiments Google’s Gemini hacked three companies.
The AI companies are investigating tens of thousands of incidents involving their agents, according to a report in Axios . These episodes have heightened fears about AI agents taking actions without human prompting.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in