GPT-6 Astra Just Went CRITICAL...
Summary
OpenAI's upcoming Astra model is the first to reach the "Critical" cybersecurity threshold under its Preparedness Framework. The Information reports Astra uses a new "looped transformers" technique allowing latent space reasoning, bypassing readable text. This architecture raises concerns about understanding model behavior, as highlighted by a recent paper from OpenAI, Anthropic, and Google DeepMind.
Why it matters
Why it matters: OpenAI's Astra model hitting a "Critical" cybersecurity threshold and employing "looped transformers" signals a shift towards AI reasoning in non-interpretable latent spaces. This technique, which bypasses readable text, echoes concerns from a joint paper by OpenAI, Anthropic, and Google DeepMind about losing visibility into AI decision-making. This development affects AI safety researchers and developers across the field, including competitors like Google Gemini and Anthropic, who are also pushing AI capabilities. Future focus will be on how these "looped transformers" are secured and whether transparency can be maintained.
Related: OpenAI: OpenAI’s next big AI model has ‘entered the AGI era’ · Google: GPT-6 Astra Just Went CRITICAL... · Microsoft: What's new in AI? · Meta: Get the full story behind the light · xAI: Sam Altman "AGI by December"
Rated high: a launch, deal or policy change that changes what you can do or what it costs.
- GPT-6 Astra Is Here—and OpenAI Thinks It May Kick Off the AGI EraWIRED AI · Web · Sep 3, 2026
- OpenAI’s new reasoning technique alarms AI safety expertsTechCrunch AI · Web · Sep 2, 2026
- Researchers fear safety disaster ahead of OpenAI’s Astra releaseThe Verge AI · Web · Sep 2, 2026
- OpenAI hails ‘new era of artificial general intelligence’ with Astra model releaseThe Guardian AI · Web · Sep 3, 2026
- The Most Overhyped and Underhyped New AI ModelsMatt Wolfe · Web · Sep 2, 2026
- OpenAI’s Altman Unveils Astra as a New Step Toward AGI | The Close 9/3/2026Bloomberg Technology · Web · Sep 3, 2026
- Fable 5.1 just smoked ASTRA...Wes Roth
- Automated researchers can reliably mitigate alignment failuresAnthropic Research
- CamoDocs: A Poisoning Attack Against Retrieval-Augmented Language Models Using Camouflaged DocumentsarXiv cs.AI
- Toy models of superpositionAnthropic Research
- Language models (mostly) know what they knowAnthropic Research
- GPT-6 Goes Rogue? The HuggingFace Incident, Sans HypeAI Explained
Questions people ask
- What is OpenAI's Preparedness Framework?
- The Preparedness Framework is a system OpenAI uses to evaluate AI models. Astra is the first model to reach the "Critical" cybersecurity threshold under this framework.
- What are "looped transformers"?
- This is a new technique reportedly used by Astra, allowing models to reason in latent space instead of readable text. It is also referred to as recurrent depth.
- Why is latent space reasoning a concern?
- A paper from OpenAI, Anthropic, and Google DeepMind warned that this type of reasoning could break the ability to understand what AI models are thinking.
More from Wes Roth 11 more
Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to Wes Roth.














