Big Pickle on SWE Atlas – Codebase QnA
A report on the performance of 'big-pickle', a stealth AI model, on the SWE Atlas Codebase QnA benchmark. The model demonstrates competitive results compared to existing leaderboard entries when using specific scaffolding.
Why it matters
Benchmarking AI models on real-world codebase tasks is critical for understanding the practical utility of LLMs in software engineering workflows.
Task Resolve Rate: 50.8% (63/124) — big-pickle , the free stealth model on OpenCode Zen, evaluated on Scale AI's SWE Atlas Codebase QnA benchmark using the mini-swe-agent scaffold.
Run on 2026-08-11 with the official open-source harness, task data, and judge model.
Against the official SWE Atlas QnA leaderboard (updated 2026-07-28):
Within the Mini-SWE-Agent scaffold class — the apples-to-apples comparison — this run outscores every entry on the official leaderboard , and it also tops the Codex-scaffold GPT entries. Only the two Claude models running on their native Claude Code scaffold score higher. Note the caveats below before treating this as a leaderboard-equivalent number.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in