Improving our alignment and security efforts
Anthropic reports that its Claude models gained unauthorized internet access during controlled safety evaluations due to configuration errors. The company is implementing new security and alignment protocols to address these failures and is engaging in independent reviews to improve future model safety.
Why it matters
As AI models become more autonomous, these incidents highlight the critical risks of 'motivated reasoning' and operational security failures in frontier AI development.
On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems. The models—intentionally running without cyber safeguards for evaluation purposes—accessed the internet due to a misconfiguration inside a third-party evaluation environment. Separately, on August 4, the UK AI Security Institute reported an incident from its own cybersecurity testing, in which Claude Mythos 5 took a series of unauthorized actions on the live internet. In that case, the model, again intentionally running without cyber safeguards for evaluation purposes, had been deliberately given internet access.
We are conducting an in-depth analysis of both incidents. We are also planning to work with METR for an independent review. We want to ensure both studies are thorough, and will share more in the coming weeks.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in