Claude Fable 5.1 made me a really nice animated pelican

Source: Simon Willison By Simon Willison
Image: Simon Willison

Today is Claude Fable (and Mythos) 5.1 day . Anthropic say that Fable 5.1 "sets a new standard for coding, knowledge work, and long-running problem-solving tasks". Their announcement spends a notable amount of time on scientific research, boasting of a 52.6% score on the brand new Terminal-Bench-Science 0.1 benchmark (first announced on August 27th ),…

Opening of the original on Simon Willison

Summary

Anthropic released Claude Fable 5.1, a new model claiming improved performance in coding, knowledge work, and problem-solving. The model achieved a 52.6% score on the new Terminal-Bench-Science 0.1 benchmark, significantly outperforming competitors like GPT-5.6 Sol. The author tested Fable 5.1's ability to generate animated pelicans across its five reasoning levels, noting that benchmark comparisons are most insightful within model families and across different reasoning efforts.

Why it matters

Why it matters: Anthropic's Fable 5.1 shows a substantial leap in scientific reasoning benchmarks, a key area for AI development. This release challenges competitors like OpenAI and Google Gemini by setting a new bar for specialized tasks. The author's focus on the 'pelican benchmark' and reasoning levels highlights the ongoing effort to evaluate AI capabilities beyond standard metrics. Future attention should be on how this scientific prowess translates to real-world applications and how other labs respond to this benchmark performance.

Read this on Simon Willison
Opens in a new tab. Subvolts summarizes and links; the full piece belongs to Simon Willison.
Where the other five stand

Related: OpenAI: Claude Fable 5.1 made me a really nice animated pelican · Anthropic: Meet Claude Fable 5.1 · Microsoft: Meet MAI-Transcribe-2: A faster and more accurate speech recognition model · Meta: Trump Administration Sides With OpenAI in New York Times Copyright Lawsuit · xAI: Sam Altman "AGI by December"

Hype check
4/5Big

Rated high: a launch, deal or policy change that changes what you can do or what it costs.

Who's talking about it
Prior coverage our earlier items on the same thing
Published
Source
Simon Willison (simonwillison.net)
Author
Simon Willison
Company
Google · Web · Trade
Products
Claude Fable 5.1, Claude Fable 5, Claude Opus 5, GPT-5.6 Sol
Summary by
Subvolts, using an AI model (how we work). Spotted a mistake? Tell us.

Questions people ask

What is new in Claude Fable 5.1?
Anthropic claims Fable 5.1 improves coding, knowledge work, and long-running problem-solving. It also achieved a high score on the new Terminal-Bench-Science 0.1 benchmark.
How does Fable 5.1 perform on science benchmarks?
Fable 5.1 scored 52.6% on the Terminal-Bench-Science 0.1 benchmark, significantly higher than previous versions and competitors like GPT-5.6 Sol.

More from Simon Willison 2 more

Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to Simon Willison.