Language models (mostly) know what they know

Source: Anthropic Research
Image: Anthropic Research

We study whether language models can evaluate the validity of their own claims and predict which questions they will be able to answer correctly. We first show that larger models are well-calibrated on diverse multiple choice and true/false questions when they are provided in the right format. Thus we can approach self-evaluation on open-ended sampling…

Opening of the original on Anthropic Research

Summary

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

Read this on Anthropic Research
Opens in a new tab. Subvolts summarizes and links; the full piece belongs to Anthropic Research.
Where the other five stand

Related: OpenAI: OpenAI’s next big AI model has ‘entered the AGI era’ · Google: GPT-6 Astra Just Went CRITICAL... · Microsoft: What's new in AI? · Meta: Get the full story behind the light · xAI: Sam Altman "AGI by December"

Hype check
2/5Worth a look

Rated low: routine. Worth knowing, not worth rearranging your day for.

Who's talking about it
Prior coverage our earlier items on the same thing
Published
Source
Anthropic Research (anthropic.com)
Company
Anthropic · Official · Research
Summary by
Subvolts, using an extract from the source (how we work). Spotted a mistake? Tell us.

Questions people ask

Where can I read the full story?
On Anthropic Research. The "Read this on Anthropic Research" link above opens the original in a new tab. Subvolts publishes a summary and analysis, never the full piece.
What does this mean for Claude?
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

More from Anthropic Research 37 more

Everything from Anthropic Research →

Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to Anthropic Research.