Claude Fable 5.1 made me a really nice animated pelican

Today is Claude Fable (and Mythos) 5.1 day . Anthropic say that Fable 5.1 "sets a new standard for coding, knowledge work, and long-running problem-solving tasks". Their announcement spends a notable amount of time on scientific research, boasting of a 52.6% score on the brand new Terminal-Bench-Science 0.1 benchmark (first announced on August 27th ),…
Opening of the original on Simon Willison
Summary
Anthropic released Claude Fable 5.1, claiming it sets a new standard for coding, knowledge work, and problem-solving. The model achieved a 52.6% score on the new Terminal-Bench-Science 0.1 benchmark, significantly outperforming competitors like GPT-5.6 Sol. The author tested Fable 5.1's "pelican" benchmark performance across five reasoning levels, noting improvements within model families.
Why it matters
Why it matters: Anthropic's Fable 5.1 launch emphasizes scientific research capabilities with a strong benchmark score, aiming to lead in complex tasks. This release positions Fable 5.1 as a direct competitor to models like OpenAI's GPT-5.6 Sol, particularly in specialized scientific domains. The focus on reasoning levels suggests Anthropic is refining model control for users. Future attention should be on how these scientific benchmark gains translate to real-world applications and if other frontier labs like Google Gemini and Meta AI respond with similar specialized advancements.
Related: Google: Introducing Gemini 3.8 Flash and 3.8 Flash Cyber · Anthropic: Meet Claude Fable 5.1 · Microsoft: Meet MAI-Transcribe-2: A faster and more accurate speech recognition model · Meta: Trump Administration Sides With OpenAI in New York Times Copyright Lawsuit · xAI: Sam Altman "AGI by December"
Rated high: a launch, deal or policy change that changes what you can do or what it costs.
- GPT‑6 AstraSimon Willison · Web · Sep 3, 2026
- GPT‑6 AstraSimon Willison · Web · Sep 3, 2026
- Claude Fable AI Is Much Stranger Than The Headlines SuggestTwo Minute Papers · Web · Sep 3, 2026
- Fable 5.1 just smoked ASTRA...Wes Roth · Web · Sep 1, 2026
- Fable 5.1 just smoked ASTRA...Wes Roth · Web · Sep 1, 2026
- OpenAI’s Altman Unveils Astra as a New Step Toward AGI | The Close 9/3/2026Bloomberg Technology · Web · Sep 3, 2026
- Fable 5.1 just smoked ASTRA...Wes Roth
- Two AI Models Set to “stir government urgency”, But Will This Challenge Undo Them?AI Explained
- Path to Astra: critical capabilities and frontier safeguardsOpenAI News
- ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide GenerationarXiv cs.AI
- Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language ModelsarXiv cs.AI
- AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic WorldsarXiv cs.AI
Questions people ask
- What are the main claims for Claude Fable 5.1?
- Anthropic states Fable 5.1 sets a new standard for coding, knowledge work, and long-running problem-solving tasks, with a notable focus on scientific research.
- How did Fable 5.1 perform on the new science benchmark?
- Fable 5.1 scored 52.6% on the Terminal-Bench-Science 0.1 benchmark, significantly higher than Fable 5 (24.7%), Opus 5 (29.0%), and GPT-5.6 Sol (22.4%).
- What reasoning levels does Fable 5.1 offer?
- Fable 5.1 provides five reasoning levels: low, medium, high, xhigh, and max. There is no option to turn off reasoning entirely.
More from Simon Willison 10 more
Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to Simon Willison.






