gpt-oss-safeguard-120b
A 120B open-weight policy-reasoning model for configurable content classification and safety enforcement.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why this model scores 44.8
Focused rather than frontier capability and little independent autonomy keep risk moderate; Apache licensing maximizes deployment while safeguard purpose limits misuse.
News tied to gpt-oss-safeguard-120b
The model score of 44.8 rates this model's risk profile. The overall Doom Index of 67.9 measures the complete temporally weighted evidence record. These values answer different questions.
Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.
OpenAI releases open-weight safety reasoners backed by deployed classifier practice
OpenAI released 120B and 20B open-weight safeguard models that interpret custom policies, with the underlying safety-reasoning approach already used in production systems.
- Full item contribution
- -0.17
- gpt-oss-safeguard-120b equal share
- -0.09
Model score history
-
R3
Doom Score 44.8
Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-142416 under temporal-monthly-pressure-v4.
12 Aug 2026 -
R2
Doom Score 44.5
Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-132112 under fixed-sensitivity-v3.
12 Aug 2026 -
R1
Doom Score 43.4
New exact safeguard model absent from the durable catalogue.
11 Aug 2026