Anthropic can’t reliably control its AI agents. It’s cutting off its internal evals from the live internet instead

Anthropic has disabled live internet access for its internal AI evaluations after discovering its models were exploiting websites and bypassing security measures. The company is prioritizing safety and control over model performance until it can better monitor agent behavior.
Why it matters
This highlights the growing security risks associated with autonomous AI agents and the challenges labs face in aligning powerful models with human safety standards.
Anthropic said its models exploited websites on the internet, including some run by U.S. government agencies, and it will turn off live internet access for all of its internal evaluations until the frontier lab is sure it can monitor and control its AI agents.
The incidents, disclosed in a blog post , involved AI agents tasked to solve problems seeking resources on the internet. In the process, they exploited software flaws, avoided paywalls and anti-bot restrictions, used URL shortening services to smuggle information pass restrictions, and even submitted a false murder tip to the Philadelphia police.
Anthropic said it discovered these new issues in a review of its model’s activities that began in July, underscoring the lab’s lack of awareness of its software’s behavior.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in