Kimi K3: second only to Fable 5 on AA-Briefcase

Moonshot AI's new Kimi K3 model has achieved the second-highest score on the AA-Briefcase benchmark, trailing only Claude Fable 5. While the model shows strong analytical capabilities, it is significantly more expensive to operate than its competitors.
Why it matters
The performance of Kimi K3 highlights the rapid advancement of agentic AI models and the increasing importance of specialized benchmarks for complex knowledge work.
Artificial Analysis K Artificial Analysis Models Coding Agents Speech, Image, Video Hardware Leaderboards About AI Trends Arenas K All articles July 21, 2026
Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging nearly an hour per task
Last week Kimi (Moonshot AI) released Kimi K3, a 2.8T parameter model that scores 57 on the Artificial Analysis Intelligence Index, comparable to models such as Opus 4.8 and GPT-5.5. On AA-Briefcase, Kimi K3 scores an Elo of 1543, a +727 improvement over Kimi K2.6 and the second highest score recorded, behind only Claude Fable 5 (1574)
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in