Translating Claude’s thoughts into language

Source: Anthropic on YouTube By Anthropic

Summary

AI models like Claude talk in words but think in numbers. These numbers, called activations, encode Claude’s thoughts, but not in a language we can read. We are introducing Natural Language Autoencoders, or NLAs, which translate AI models’ activations into readable text. NLAs have already helped us improve how we test our models for safety and better understand why they do what they do. Read more about this research on our blog: https://www.anthropic.com/research/natural-language-autoencoders

Read this on Anthropic on YouTube
Opens in a new tab. Subvolts summarizes and links; the full piece belongs to Anthropic on YouTube.
Where the other five stand

Related: OpenAI: Two Rival Bets on AGI: Google I/O Highlights · Google: Two Rival Bets on AGI: Google I/O Highlights

Hype check
2/5Worth a look

Rated low: routine. Worth knowing, not worth rearranging your day for.

Prior coverage our earlier items on the same thing
Published
Source
Anthropic on YouTube (youtube.com)
Author
Anthropic
Company
Anthropic · Official · Videos
Summary by
Subvolts, using an extract from the source (how we work). Spotted a mistake? Tell us.

Questions people ask

Where can I read the full video?
On Anthropic on YouTube. The "Read this on Anthropic on YouTube" link above opens the original in a new tab. Subvolts publishes a summary and analysis, never the full piece.
What does this mean for Claude?
AI models like Claude talk in words but think in numbers. These numbers, called activations, encode Claude’s thoughts, but not…

More from Anthropic on YouTube 14 more

Everything from Anthropic on YouTube →

Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to Anthropic on YouTube.