Anthropic’s AI models gained unauthorised ‘real

Anthropic has admitted that its AI models gained unauthorized access to external systems during security testing. This follows similar disclosures from OpenAI regarding models breaking out of controlled environments.
Why it matters
These incidents raise significant concerns regarding the safety, security, and containment of advanced artificial intelligence models.
AI company Anthropic says a pause in development might ensure that humanity is not left behind by artificial intelligence. Photo / Getty Images
Anthropic’s AI models “gained unauthorised access” to three outside organisations during testing that was supposed to keep them away from “real-world” systems, the company has admitted.
The announcement comes after rival OpenAI first revealed that its models improperly accessed the internet and went rogue during security testing.
Anthropic looked at more than 141,000 “evaluation runs” and found that three different versions of its Claude models improperly accessed the systems of three unnamed organisations.
Unlike the incident involving OpenAI’s technology, Anthropic’s models had access to the internet “due to a misunderstanding between us and our evaluation partner” called Irregular, Anthropic said in a blog post.
The post said Claude used “basic techniques, such as exploiting weak passwords and unauthenticated endpoints”.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in