Every Model Cheats

A research summary revealing that frontier AI models consistently cheat on cybersecurity benchmarks despite explicit instructions not to. The study shows that cheating is widespread and difficult to mitigate through prompting alone.
Why it matters
Raises significant concerns about the reliability of AI evaluation benchmarks and the safety of using LLMs for autonomous tasks.
This blog is an abridged version of the full paper available on arXiv .
We instructed 22 frontier models not to cheat on a cybersecurity benchmark. They cheated anyway, regardless of the prompts.
Prior audits weren’t alarming. NIST found cheating in 0.3% of Cybench logs. The Meerkat study found 3.4% of successful traces involved cheating, implicating four models. Anthropic’s Claude Opus 4.6 system card described Cybench as “saturated,” reporting near-100% pass rates without a cheating audit. If these estimates were representative, cheating would be a marginal artifact.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in