13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS

A status update on an AI leaderboard tracking the performance of various models and agents across software engineering tasks. The update details the addition of new models and the deprecation of older versions.
Why it matters
It tracks the rapid iteration cycle of LLMs, providing benchmarks for developers choosing models for coding tasks.
0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36
111 problems from 65 repositories selected within the current time window.
gemini-2.5-flash-preview-05-20 no-thinking
Added new models to the leaderboad: GLM 5.2, DeepSeek-V4 Pro, DeepSeek-V4 Flash, MiMo V2.5 Pro, Qwen3.6-35B-A3B, Qwen3.6-27B and Gemma 4 31B.
Added new models to the leaderboad: Gemini 3.5 Flash and MiniMax M3.
Added new models to the leaderboad: Claude Opus 4.8.
Added new models to the leaderboad: gpt-5.5-2026-04-23-xhigh, gpt-5.5-2026-04-23-medium, gpt-5.4-2026-03-05-medium, Claude Opus 4.7, and Kimi K2.6.
Re-run the Junie with Claude Opus 4.6 as the primary model.
Added new models to the leaderboad: GLM-5.1, Qwen3.5-27B, Cursor, Gemma 4 31B and MiniMax M2.7.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in