OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings

OpenAI claims its GPT-5.6 Sol model outperforms Anthropic's Claude Opus 5 on the ARC-AGI-3 benchmark using specific API settings. The results have sparked debate regarding testing standards and the role of technical setups in AI performance.
Why it matters
Standardized benchmarking is critical for evaluating the true capabilities of competing AI models in a rapidly evolving field.
Update AI in practice Copy the url to clipboard Share this article Go to comment section OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 30, 2026 Nano Banana Pro prompted by THE DECODER Ask about this article… Search Update – Jul 30, 2026 Added ARC Prize statements Update:
ARC Prize co-founder François Chollet responded to OpenAI's results by distinguishing between two kinds of test setups. Harnesses "custom-made to solve the benchmark or that contain knowledge about the benchmark format" are off limits, he said. General-purpose API settings "that were not developed for ARC-AGI-3 and that are available to all API users" are fair game. In effect, Chollet is conceding that ARC Prize's own GPT-5.6 Sol score put OpenAI at a disadvantage.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in