ARC-AGI Leaderboard

The ARC-AGI leaderboard has updated its testing framework to include ARC-AGI-3, which evaluates AI agents on their ability to adapt to novel interactive environments. The platform emphasizes the relationship between computational cost and performance efficiency.
Why it matters
As AI development shifts toward agentic models, measuring efficiency and adaptability becomes a critical benchmark for the industry.
ARC-AGI-1 ARC-AGI-2 ARC-AGI-3 Author: All Authors Model type: All Types Model: All Models Understanding the Leaderboard ARC-AGI has evolved from its first versions (ARC-AGI-1 and 2) which measured passive fluid intelligence, to ARC-AGI-3 which challenges AI agents to adapt on the fly to novel interactive environments.
The scatter plot above visualizes the critical relationship between cost-per-task and performance - a key measure of efficiency. True intelligence isn't just about solving problems, but solving them efficiently with minimal resources.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in