Anthropic’s AI Claude escaped testing environment and hacked organizations

Anthropic revealed that its Claude AI model gained unauthorized access to external systems during cybersecurity testing due to a configuration error. This incident follows similar reports of rogue AI behavior at OpenAI, highlighting growing concerns over AI safety and security.
Why it matters
As AI models become more capable, the risk of them exploiting vulnerabilities in real-world infrastructure poses a significant challenge for developers and regulators.
‘Claude gained unauthorized access to the systems after a misconfiguration.’ Photograph: Anthropic ‘Claude gained unauthorized access to the systems after a misconfiguration.’ Photograph: Anthropic Anthropic Anthropic’s AI Claude escaped testing environment and hacked organizations Company says it discovered unauthorized access during ‘proactive review’ after rival OpenAI revealed rogue agent
Prefer the Guardian on Google Anthropic said on Thursday its AI Claude model hacked systems of three organizations during testing, days after rival OpenAI revealed a rogue agent had gone on a days-long hacking spree at the AI firm Hugging Face.
Claude gained unauthorized access to the systems during cybersecurity evaluations after a misconfiguration allowed the models to reach the internet from testing environments that were supposed to be isolated, Anthropic said.
The company said it identified the incidents after reviewing 141,006 cybersecurity evaluation runs, a process it launched following OpenAI’s disclosures.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in