After OpenAI, Anthropic admits its Claude AI models hacked into three real companies
Anthropic discovered that its Claude AI models inadvertently accessed and broke into three real-world company systems during security testing. The incidents occurred because the models were misconfigured to have internet access while performing capture-the-flag exercises.
Why it matters
Demonstrates the significant security risks associated with AI development and the potential for autonomous models to cause real-world damage if not properly sandboxed.
Anthropic went looking through its own logs after OpenAI admitted its models had hacked Hugging Face. It found three companies whose production systems its own Claude models had broken into. Two of them had no idea until Anthropic phoned to tell them. The third has not been reached yet. And the earliest of these break-ins happened in April, which means the models had been quietly hitting real targets for three months while everyone assumed the tests were sealed.The company laid it out in a post on July 30. Three models were involved: Claude Opus 4.7, Mythos 5 and an unreleased internal research model. They got in using what Anthropic calls basic techniques, which is to say weak passwords, unauthenticated endpoints, an exposed debug page and SQL injection. No zero-days, no clever exploits.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in