OpenAI and Anthropic publish joint model safety evaluations
OpenAI and Anthropic jointly evaluated each other's frontier models across alignment, misuse and capability tests, publishing strengths, failures and limits of extrapolating the results to real-world behavior.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
Cross-laboratory evaluation exposed concrete safety strengths and failures across exact frontier checkpoints, improving independent scrutiny while remaining evaluation evidence rather than a deployed control.
Assessment history
-
R1
Away 28 · confidence 82
New August 2025 cross-lab safety evaluation with exact model relationships and no durable collision.
12 Aug 2026
Share this page
-
DoomBench assesses “OpenAI and Anthropic publish joint model safety evaluations” as evidence moving away from doom, with magnitude 28 and confidence 82 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “OpenAI and Anthropic publish joint model safety evaluations” is based on reporting from OpenAI and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “OpenAI and Anthropic publish joint model safety evaluations” as follows: OpenAI and Anthropic jointly evaluated each other's frontier models across alignment, misuse and capability tests, publishing strengths,...
https://www.doombench.com/news/openai-and-anthropic-publish-joint-model-safety-evaluations-2025-08-27