Article may be outdated

This article is 66 days old. Some details may have changed since publication.

Hacker News·4 min read·medium

Benchmarking Opus 5 on SlopCodeBench

D
dhorthy
Benchmarking Opus 5 on SlopCodeBench
✦AI Summary

A developer evaluates the performance of advanced AI models like Claude Opus 5 on SlopCodeBench, a new benchmark designed to test long-horizon coding capabilities. The results suggest that even top-tier models struggle with real-world software engineering tasks that require iterative development without human intervention.

Why it matters

It highlights the current limitations of AI in autonomous software development, suggesting that models still lack the reliability needed for complex, multi-step coding projects.

✦Dive DeeperCreate a free account to unlock

I've written before something along the lines of:

That wasn't entirely true. I love nothing more than burying a good lede.

Last Friday I dug into SlopCodeBench , a new-ish (March 2026) long-horizon coding benchmark from @GOrlanski 's lab at UW Madison. It addresses the thing that bothers me most about coding benchmarks - that even "larger" more complex benchmarks still divulge the whole problem up front:

In contrast, each challenge in SlopCodeBench has multiple "checkpoints" - the model doesn't know the whole problem up front, it has to evolve the codebase over time as new requirements are divulged.

It's a good paper. It's not that long. You should read it .

What's cool about this benchmark is that it is unsaturated - at the time of running, the best models available, GPT-5.4 and Opus 4.6, got 11% and 17% strict pass rates, respectively.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologybusiness
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in