OpenAI · Official · Blog Article
AI-written critiques help humans notice flaws
Summary
We trained “critique-writing” models to describe flaws in summaries. Human evaluators find flaws in summaries much more often when shown our model’s critiques. Larger models are better at self-critiquing, with scale improving critique-writing more than summary-writing. This shows promise for using AI systems to assist human supervision of AI systems on difficult tasks.
Opens in a new tab. Subvolts summarizes and links; the full piece belongs to OpenAI News.
Hype check
2/5Worth a look
Rated low: routine. Worth knowing, not worth rearranging your day for.
Prior coverage our earlier items on the same thing
- Best practices for deploying language modelsOpenAI News
- Teaching models to express their uncertainty in wordsOpenAI News
- Lessons learned on language model safety and misuseOpenAI News
- A research agenda for assessing the economic impacts of code generation modelsOpenAI News
- Economic impacts research at OpenAIOpenAI News
- Aligning language models to follow instructionsOpenAI News
Questions people ask
- Where can I read the full story?
- On OpenAI News. The "Read this on OpenAI News" link above opens the original in a new tab. Subvolts publishes a summary and analysis, never the full piece.
- What does this mean for ChatGPT?
- We trained “critique-writing” models to describe flaws in summaries. Human evaluators find flaws in summaries much more often when shown…
More from OpenAI News 1017 more
Page generated Sep 3, 2026. Summaries are Subvolts' own; the story belongs to OpenAI News.


