OpenAI reports deliberative alignment in deployed o-series models
OpenAI reported that training reasoning models to apply written safety policies improved refusal and jailbreak performance, with the method used in the deployed o1 family.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The original dated safety result is paired with separate first-party deployment evidence in OpenAI's December 5 ChatGPT Pro and December 17 API releases. The method improved policy reasoning but did not eliminate jailbreaks or dependence on supplied specifications.
Assessment history
-
R1
Away 54 · confidence 96
New December 2024 safety research with separate dated practical deployment evidence and no durable collision.
12 Aug 2026
Share this page
-
DoomBench assesses “OpenAI reports deliberative alignment in deployed o-series models” as evidence moving away from doom, with magnitude 54 and confidence 96 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “OpenAI reports deliberative alignment in deployed o-series models” is based on reporting from OpenAI and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “OpenAI reports deliberative alignment in deployed o-series models” as follows: OpenAI reported that training reasoning models to apply written safety policies improved refusal and jailbreak performance, with the...
https://www.doombench.com/news/openai-reports-deliberative-alignment-in-deployed-o-series-models-2024-12-20