gpt-oss-safeguard technical report
Summary
gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are two open-weight reasoning models post-trained from the gpt-oss models and trained to reason from a provided policy in order to label content under that policy. In this report, we describe gpt-oss-safeguard’s capabilities and provide our baseline safety evaluations on the gpt-oss-safeguard models, using the underlying gpt-oss models as a baseline. For more information about the development and architecture of the underlying gpt-oss models, see the original gpt-oss model model card.
Related: Google: Mapping, modeling, and understanding nature with AI · Meta: Introducing the Segment Anything Playground | AI at Meta
Rated low: routine. Worth knowing, not worth rearranging your day for.
- SafetyKit scales risk agents with OpenAI’s most capable modelsOpenAI News
- Why language models hallucinateOpenAI News
- OpenAI and Anthropic share findings from a joint safety evaluationOpenAI News
- gpt-oss-120b & gpt-oss-20b Model CardOpenAI News
- Introducing HealthBenchOpenAI News
- Deliberative alignment: reasoning enables safer language modelsOpenAI News
Questions people ask
- Where can I read the full story?
- On OpenAI News. The "Read this on OpenAI News" link above opens the original in a new tab. Subvolts publishes a summary and analysis, never the full piece.
- What does this mean for ChatGPT?
- gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are two open-weight reasoning models post-trained from the gpt-oss models and trained to reason from a provided…
More from OpenAI News 1017 more
Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to OpenAI News.


