Safety and alignment

OpenAI and Anthropic publish joint model safety evaluations

OpenAI and Anthropic jointly evaluated each other's frontier models across alignment, misuse and capability tests, publishing strengths, failures and limits of extrapolating the results to real-world behavior.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM28confidence 82/100

Why it moved the index

Cross-laboratory evaluation exposed concrete safety strengths and failures across exact frontier checkpoints, improving independent scrutiny while remaining evaluation evidence rather than a deployed control.

AUDIT TRAIL

Assessment history

  1. R1
    Away 28 · confidence 82

    New August 2025 cross-lab safety evaluation with exact model relationships and no durable collision.

    12 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for OpenAI and Anthropic publish joint model safety evaluations.
  1. DoomBench assesses “OpenAI and Anthropic publish joint model safety evaluations” as evidence moving away from doom, with magnitude 28 and confidence 82 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “OpenAI and Anthropic publish joint model safety evaluations” is based on reporting from OpenAI and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “OpenAI and Anthropic publish joint model safety evaluations” as follows: OpenAI and Anthropic jointly evaluated each other's frontier models across alignment, misuse and capability tests, publishing strengths,...

    https://www.doombench.com/news/openai-and-anthropic-publish-joint-model-safety-evaluations-2025-08-27