Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

Cognition has released its new SWE-2 AI model, which shows significant performance gains on coding benchmarks like Terminal-Bench 2.1. While the model demonstrates high efficiency in specific tasks, it still lags behind frontier models in long-horizon agentic capabilities.
Why it matters
The rapid evolution of AI coding agents is transforming software development workflows and competitive benchmarks in the tech industry.
2.8T total params, 104B active per token (MoE) - the Kimi K3 base with Cognition’s post-training on top, and the first time Cognition has scaled RL into the multi-trillion-parameter regime. The base had already been RL-heavy for agentic coding; Cognition’s pass added another 5 to 6 points on most benchmarks.
Benchmarks (Cognition self-reported): FrontierCode 1.1 Main 50.0, DeepSWE 1.1 73.0, Terminal-Bench 2.1 92.8, Terminal-Bench 4.0 27.3. The headline: 50.0 on FrontierCode is one point behind Claude Fable 5.1 (50.9) and 3.3 behind GPT-6 Astra (53.3) - at a claimed 64% lower cost than Fable 5.1 and a quarter of Astra’s. Terminal-Bench 2.1 is the highest number in the published table. The soft spot is Terminal-Bench 4.0, where SWE-2’s 27.3 trails Fable 5.1 (55.8) and GPT-6 Astra (57.9) by a wide margin - long-horizon agentic work is where the gap to the frontier still lives.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in