Breaking Claude Code Opus 5 Auto Mode

A security researcher demonstrates that Claude Code's 'Auto Mode' is vulnerable to indirect prompt injection attacks, achieving high success rates in hijacking the agent. This finding challenges recent third-party safety evaluations that claimed zero attack success for the model.
Why it matters
It highlights the risks of relying on automated AI agents for sensitive tasks and suggests that current safety classifiers may be insufficient against sophisticated prompt injection.
In this post, we explore how a simple website summary request hijacks Claude Code Opus 5 in Auto Mode and achieves code execution with 60-80% attack success rate using a small sample size.
Reports on a security vulnerability and challenges a corporate claim with empirical testing.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in