Measuring the persuasiveness of language models

While people have long questioned whether AI models may, at some point, become as persuasive as humans in changing people's minds, there has been limited empirical research into the relationship between model scale and the degree of persuasiveness across model outputs. To address this, we developed a basic method to measure persuasiveness, and used it…
Opening of the original on Anthropic Research
Summary
Anthropic developed a way to test how persuasive language models (LMs) are, and analyzed how persuasiveness scales across different versions of Claude.
Related: OpenAI: Introducing GPT-6 Astra: the most intelligent and aligned model in the world. · Google: Run Ray on TPU, Part 2: Ray AI libraries · Microsoft: Meet MAI-Transcribe-2: A faster and more accurate speech recognition model · Meta: MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet · xAI: Sam Altman "AGI by December"
Rated low: routine. Worth knowing, not worth rearranging your day for.
- Four major AI models suffer rare overlapping downtimeArs Technica AI · Web · Sep 3, 2026
- Claude Fable 5.1 made me a really nice animated pelicanSimon Willison · Web · Sep 1, 2026
- The mystery is solved... and the answer is 40x cheaper than ClaudeFireship · Web · Sep 1, 2026
- Claude Fable AI Is Much Stranger Than The Headlines SuggestTwo Minute Papers · Web · Sep 3, 2026
- The Most Overhyped and Underhyped New AI ModelsMatt Wolfe · Web · Sep 2, 2026
- [AINews] Claude Fable/Mythos 5.1: new SOTA model, 75% cache price cut but 70% more output tokensLatent Space · Web · Sep 2, 2026
- Toy models of superpositionAnthropic Research
- Language models (mostly) know what they knowAnthropic Research
Questions people ask
- Where can I read the full story?
- On Anthropic Research. The "Read this on Anthropic Research" link above opens the original in a new tab. Subvolts publishes a summary and analysis, never the full piece.
- What does this mean for Claude?
- Anthropic developed a way to test how persuasive language models (LMs) are, and analyzed how persuasiveness scales across different versions…
More from Anthropic Research 37 more
Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to Anthropic Research.



