Safety and alignment

OpenAI reports deliberative alignment in deployed o-series models

OpenAI reported that training reasoning models to apply written safety policies improved refusal and jailbreak performance, with the method used in the deployed o1 family.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM54confidence 96/100

Why it moved the index

The original dated safety result is paired with separate first-party deployment evidence in OpenAI's December 5 ChatGPT Pro and December 17 API releases. The method improved policy reasoning but did not eliminate jailbreaks or dependence on supplied specifications.

AUDIT TRAIL

Assessment history

  1. R1
    Away 54 · confidence 96

    New December 2024 safety research with separate dated practical deployment evidence and no durable collision.

    12 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for OpenAI reports deliberative alignment in deployed o-series models.
  1. DoomBench assesses “OpenAI reports deliberative alignment in deployed o-series models” as evidence moving away from doom, with magnitude 54 and confidence 96 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “OpenAI reports deliberative alignment in deployed o-series models” is based on reporting from OpenAI and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “OpenAI reports deliberative alignment in deployed o-series models” as follows: OpenAI reported that training reasoning models to apply written safety policies improved refusal and jailbreak performance, with the...

    https://www.doombench.com/news/openai-reports-deliberative-alignment-in-deployed-o-series-models-2024-12-20