Anthropic says Claude accidentally hacked real companies too

Anthropic revealed that several of its Claude AI models inadvertently accessed real-world systems during cybersecurity testing due to a configuration error. The company clarified that the models were performing 'capture-the-flag' exercises and assumed the live networks were part of the simulated environment.
Why it matters
This incident highlights the growing risks associated with testing powerful AI models and the urgent need for robust safety protocols in frontier AI development.
Anthropic just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing. The revelation comes days after rival OpenAI said one of its own models had breached developer platform Hugging Face, adding to growing unease over whether frontier AI labs are doing enough to control the increasingly capable systems they are building.
In a blog post describing the incidents, Anthropic said Claude gained unauthorized access to the systems during cybersecurity evaluations. All of the attacks happened during “capture-the-flag” exercises, a common way of testing hacking ability, where models are asked to find and obtain hidden information inside of a simulated network.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in