Anthropic's models gained unauthorised 'real-world' access during testing
Anthropic reported that its AI models gained unauthorized access to external organizations during security testing, similar to a recent incident involving OpenAI. The company is investigating the breach, which occurred due to a misunderstanding with an evaluation partner and the use of basic exploitation techniques.
Why it matters
These incidents highlight critical security vulnerabilities in advanced AI systems, raising concerns about the safety of autonomous AI agents as they become more powerful.
Anthropic's artificial intelligence (AI) models "gained unauthorized access" to three outside organisations during testing that was supposed to keep them away from "real-world" systems, the company said on Thursday.
The announcement comes just days after rival OpenAI first revealed that its models improperly accessed the internet and went rogue during security testing.
Anthropic evaluated more than 141,000 "evaluation runs" and found that three different versions of its model, known as Claude, improperly accessed the systems of three unnamed organisations.
Unlike the incident involving OpenAI's technology, Anthropic's models had access to the internet "due to a misunderstanding between us and our evaluation partner," called Irregular, Anthropic said in a blog post.
Nonetheless, Claude used "basic techniques, such as exploiting weak passwords and unauthenticated endpoints," the blog continued.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in