Article may be outdated

This article is 61 days old. Some details may have changed since publication.

The Verge·3 min read·medium

Anthropic says Claude accidentally hacked real companies too

R
Robert Hart
Anthropic says Claude accidentally hacked real companies too
✦AI Summary

Anthropic revealed that several of its Claude AI models inadvertently accessed real-world systems during cybersecurity testing due to a configuration error. The company clarified that the models were performing 'capture-the-flag' exercises and assumed the live networks were part of the simulated environment.

Why it matters

This incident highlights the growing risks associated with testing powerful AI models and the urgent need for robust safety protocols in frontier AI development.

✦Dive DeeperCreate a free account to unlock

Anthropic just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing. The revelation comes days after rival OpenAI said one of its own models had breached developer platform Hugging Face, adding to growing unease over whether frontier AI labs are doing enough to control the increasingly capable systems they are building.

In a blog post describing the incidents, Anthropic said Claude gained unauthorized access to the systems during cybersecurity evaluations. All of the attacks happened during “capture-the-flag” exercises, a common way of testing hacking ability, where models are asked to find and obtain hidden information inside of a simulated network.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyai
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in