OpenAI · Official · Blog Article
OpenAI and Anthropic share findings from a joint safety evaluation
Summary
OpenAI and Anthropic share findings from a first-of-its-kind joint safety evaluation, testing each other’s models for misalignment, instruction following, hallucinations, jailbreaking, and more—highlighting progress, challenges, and the value of cross-lab collaboration.
Opens in a new tab. Subvolts summarizes and links; the full piece belongs to OpenAI News.
Where the other five stand
Related: Meta: Introducing DINOv3: Self-supervised learning for vision at unprecedented scale
Hype check
2/5Worth a look
Rated low: routine. Worth knowing, not worth rearranging your day for.
Prior coverage our earlier items on the same thing
- Introducing HealthBenchOpenAI News
- Deliberative alignment: reasoning enables safer language modelsOpenAI News
- Improving Model Safety Behavior with Rule-Based RewardsOpenAI News
- OpenAI and Los Alamos National Laboratory announce research partnershipOpenAI News
- OpenAI Red Teaming NetworkOpenAI News
- ChatGPT pluginsOpenAI News
Questions people ask
- Where can I read the full story?
- On OpenAI News. The "Read this on OpenAI News" link above opens the original in a new tab. Subvolts publishes a summary and analysis, never the full piece.
- What does this mean for ChatGPT?
- OpenAI and Anthropic share findings from a first-of-its-kind joint safety evaluation, testing each other’s models for misalignment, instruction following, hallucinations,…
More from OpenAI News 1017 more
Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to OpenAI News.



