Benchmarking Opus 5 on SlopCodeBench
A developer evaluates the performance of advanced AI models like Claude Opus 5 on SlopCodeBench, a new benchmark designed to test long-horizon coding capabilities. The results suggest that even top-tier models struggle with real-world software engineering tasks that require iterative development without human intervention.
Why it matters
It highlights the current limitations of AI in autonomous software development, suggesting that models still lack the reliability needed for complex, multi-step coding projects.
I've written before something along the lines of:
That wasn't entirely true. I love nothing more than burying a good lede.
Last Friday I dug into SlopCodeBench , a new-ish (March 2026) long-horizon coding benchmark from @GOrlanski 's lab at UW Madison. It addresses the thing that bothers me most about coding benchmarks - that even "larger" more complex benchmarks still divulge the whole problem up front:
In contrast, each challenge in SlopCodeBench has multiple "checkpoints" - the model doesn't know the whole problem up front, it has to evolve the codebase over time as new requirements are divulged.
It's a good paper. It's not that long. You should read it .
What's cool about this benchmark is that it is unsaturated - at the time of running, the best models available, GPT-5.4 and Opus 4.6, got 11% and 17% strict pass rates, respectively.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in