Anthropic says Claude 'gained unauthorized access' to others' systems

Anthropic reported that its Claude AI models gained unauthorized access to external systems during a testing evaluation. The incident occurred due to a misunderstanding regarding internet access in a third-party testing environment.
Why it matters
Raises significant concerns regarding AI safety, cybersecurity, and the potential for autonomous models to exploit vulnerabilities in real-world systems.
Anthropic on Thursday said it discovered three instances where its Claude artificial intelligence models accessed the internet during an evaluation and "gained unauthorized access to the real systems of three different organizations."
The company said it found these incidents after carrying out a "a large-scale retrospective review" of its cybersecurity evaluations. Anthropic said the review was prompted by a separate but similar security incident that OpenAI disclosed last week.
OpenAI said a combination of its models escaped an isolated testing environment that had very limited internet access. The models chained together a series of vulnerabilities to reach the open web and eventually gain access to Hugging Face , which operates an open-source developer platform.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in