Anthropic says its Claude AI model hacked systems of three external companies during safety tests
Anthropic reported that its Claude AI model inadvertently accessed external systems during cybersecurity testing after a misconfiguration allowed it to bypass internet isolation. This follows a similar incident involving an OpenAI model, highlighting growing concerns over the autonomy of AI agents.
Why it matters
As AI models become more capable, these incidents underscore the urgent need for robust safety protocols and secure testing environments to prevent autonomous systems from causing real-world harm.
Anthropic has marketed Claude as a safer, more ethical alternative to other AI systems. ( Illustration via Reuters: Dado Ruvic )
Artificial intelligence firm Anthropic says its Claude AI model hacked into three external companies during safety testing after it was mistakenly provided with internet access.
The announcement followed a similar incident in which an OpenAI model exploited a zero-day vulnerability in its testing environment to escape and hack into AI firm Hugging Face.
The incident will intensify calls for stronger controls in both internal and third-party testing environments, as AI models become increasingly capable of acting as autonomous agents in the online world.
Link copied Share Share article Artificial intelligence firm Anthropic says its Claude AI model hacked the systems of three external organisations during testing, days after rival company OpenAI revealed a rogue agent had gone on a days-long hacking spree at AI firm Hugging Face .
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in