Hacker News·4 min read·hard

OpenAI's GPT-6 Astra on ARC-AGI-3

V
vignesh_warar
OpenAI's GPT-6 Astra on ARC-AGI-3
AI Summary

OpenAI's GPT-6 Astra has achieved state-of-the-art results on the ARC-AGI-3 benchmark, demonstrating advanced agentic capabilities. The model successfully used symbolic world models to solve complex, abstract environments, often outperforming human baselines in action efficiency.

Why it matters

Success on the ARC-AGI benchmark is a key indicator of progress toward Artificial General Intelligence, specifically in autonomous reasoning and planning.

Dive DeeperCreate a free account to unlock

Summary GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment. , and 99.9% for $19K with a Provider Adapter harness The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work. . GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels. A key behavior observed in GPT-6 Astra was its ability to turn unfamiliar environments into compact symbolic world models. It represented game mechanics as logical rules and developed its own domain-specific language shorthand to track state and plan actions. ARC-AGI-3 ARC-AGI-3 is a benchmark for studying agentic intelligence through novel, abstract, turn-based environments.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →

Also covering this story

7 other newsrooms covered this event. We read each version separately.

technologyaiscience

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in