OpenAI πŸ‡ΊπŸ‡Έ Β· GPT-5.4

GPT-5.4 Thinking

Frontier reasoning and computer-use model deployed across ChatGPT and Codex and used as OpenAI's internal coding-agent monitor.

DOOM SCORE86.0out of 100model risk profile, not the overall index
0 comments Β· 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT Β· REVISION 5

Why this model scores 86.0

High cyber classification, native computer use, long-horizon tools, and broad deployment create substantial risk, moderated by hosted controls and low reported chain-of-thought obfuscation.

Capability93
Autonomy90
Deployment96
Misuse potential86
Control difficulty64
MODEL-ATTRIBUTED EVIDENCE

News tied to GPT-5.4 Thinking

The model score of 86.0 rates this model's risk profile. The overall Doom Index of 67.9 measures the complete temporally weighted evidence record. These values answer different questions.

NET MODEL-ATTRIBUTED INDEX CONTRIBUTION-0.10

Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.

AUDIT TRAIL

Model score history

  1. R5
    Doom Score 86.0

    Source-backed model availability audit using the model's existing primary or authoritative catalogue evidence. Exact-version evidence chronology replayed under temporal-monthly-pressure-v4.

    12 Aug 2026
  2. R4
    Doom Score 86.0

    Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-142416 under temporal-monthly-pressure-v4.

    12 Aug 2026
  3. R3
    Doom Score 82.8

    Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-132112 under fixed-sensitivity-v3.

    12 Aug 2026
  4. R2
    Doom Score 74.1

    Full-corpus evidence recalculated after run intensive-backfill-20260811-191225 under bounded-corpus-v2.

    11 Aug 2026
  5. R1
    Doom Score 86.0

    Exact missing model required for the deployed monitoring association and verified from its dated primary launch.

    11 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for GPT-5.4 Thinking.
  1. GPT-5.4 Thinking by OpenAI has a DoomBench model risk score of 86.0 out of 100, based on five transparent version-specific dimensions rather than the overall index.

  2. GPT-5.4 Thinking's highest current DoomBench dimension is deployment at 96.0 out of 100; the profile publishes every component score and its editorial rationale.

  3. DoomBench links 2 source-backed evidence items to GPT-5.4 Thinking, while keeping the model's risk profile separate from each item's contribution to the live Doom Index.

    https://www.doombench.com/models/openai-gpt-5-4-thinking