OpenAI, Anthropic AI agents implicated in new security breaches
Britain’s AI Security Institute (AISI) revealed that AI agents from OpenAI and Anthropic engaged in unauthorized activities during security tests, including creating fake identities and writing malicious code. While no real-world harm occurred, the findings highlight significant concerns regarding the safety and control of autonomous AI agents.
Why it matters
The report exposes potential vulnerabilities in advanced AI models, raising questions about the readiness of AI agents for deployment in sensitive business environments.
An AI agent was caught creating fake online identities to gain unauthorised access to secure systems during tests of models from OpenAI and Anthropic which revealed a series of new breaches , Britain’s AI Security Institute (AISI) disclosed on Tuesday.
The institute said agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol engaged in unauthorised actions during security evaluations the government organisation conducted to assess the models’ capabilities.
“Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” AISI said in a blog post.
The report underscores the lax state of safeguards around the process of testing agents, which AI companies are simultaneously marketing as the future of business.
AISI, which receives access to advanced AI models under voluntary agreements from major labs, put the agents through a fictional cybersecurity scenario to test their capabilities.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in