Claude Fable 5.1 made me a nice animated pelican

Anthropic has released Claude Fable 5.1, which claims significant improvements in scientific reasoning benchmarks. A user analysis suggests the model's reasoning effort levels behave inconsistently, with some settings skipping reasoning traces entirely for simple tasks.
Why it matters
As AI models integrate complex reasoning capabilities, understanding how these features function in practice is critical for developers and researchers relying on them for high-stakes tasks.
Today is Claude Fable (and Mythos) 5.1 day . Anthropic say that Fable 5.1 “sets a new standard for coding, knowledge work, and long-running problem-solving tasks”. Their announcement spends a notable amount of time on scientific research, boasting of a 52.6% score on the brand new Terminal-Bench-Science 0.1 benchmark (first announced on August 27th ), up from 24.7% for Fable 5, 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol. Other benchmarks show slightly improved scores, but none as impressive as the Science one.
Back in July I wrote about how I was losing faith in the pelican benchmark—its connection to how good the models were at other tasks didn’t seem to hold as strongly as it did back in 2025 . The most interesting insights I get from it now are comparisons within model families, and particularly comparisons for the same prompt at different reasoning effort levels.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in