Claude Opus 4.8
A highly capable agentic model with stronger coding, computer use, long-running work, and documented improvements in alignment behavior.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why this model scores 83.9
Near-frontier autonomy and dangerous-domain capability create substantial risk, partly offset by lower measured misalignment and deployed cyber safeguards.
News tied to Claude Opus 4.8
The model score of 83.9 rates this model's risk profile. The overall Doom Index of 67.9 measures the complete temporally weighted evidence record. These values answer different questions.
Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.
Anthropic shows automated researchers can mitigate ten alignment failures
In a controlled study, Claude autonomously developed post-training methods that improved all ten tested alignment-failure categories without measured capability loss, generalized to withheld evaluations and larger models, and closed 65% of an early Claude Opus 4.8 checkpoint's measured safety gap in 60 hours.
- Full item contribution
- -0.11
- Claude Opus 4.8 equal share
- -0.06
Anthropic opens Model Hardware Standard preview for AI-controlled laboratory equipment
Anthropic opened a limited research preview of a model-agnostic hardware interface after controlled laboratory proofs showed AI agents coordinating physical instruments, adjusting parameters, recovering from some errors, and running closed-loop experiments. One Carnegie Mellon demonstration used Claude Opus 4.8 to control incompatible lab equipment and complete a corrected dose-response run without human input.
- Full item contribution
- +0.12
- Claude Opus 4.8 equal share
- +0.12
Anthropic uses Fable 5 and Opus 4.8 for million-line code migrations
Anthropic reported that individual developers migrated ten large code packages with Fable 5, Opus 4.8 and agentic workflows. One effort produced a million lines of Rust in under two weeks with the full existing test suite passing before merge; another converted a codebase to 165,000 lines of TypeScript over a weekend using hundreds of agents and staged adversarial review.
- Full item contribution
- +0.16
- Claude Opus 4.8 equal share
- +0.08
Claude Opus 4.8 extends long-running and multi-agent work
Anthropic released Claude Opus 4.8 globally with stronger agentic and computer-use performance plus dynamic workflows capable of coordinating hundreds of parallel subagents.
- Full item contribution
- +0.38
- Claude Opus 4.8 equal share
- +0.38
Model score history
-
R6
Doom Score 83.9
Exact-version evidence chronology replayed after run doombench-20260829-065948-current-and-people-backfill under temporal-monthly-pressure-v4.
29 Aug 2026 -
R5
Doom Score 84.0
Exact-version evidence chronology replayed after run doombench-hourly-news-20260828-201012 under temporal-monthly-pressure-v4.
28 Aug 2026 -
R4
Doom Score 83.9
Source-backed model availability audit using the model's existing primary or authoritative catalogue evidence. Exact-version evidence chronology replayed under temporal-monthly-pressure-v4.
12 Aug 2026 -
R3
Doom Score 83.9
Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-132112 under fixed-sensitivity-v3.
12 Aug 2026 -
R2
Doom Score 84.1
Full-corpus evidence recalculated after run intensive-backfill-20260811-191225 under bounded-corpus-v2.
11 Aug 2026 -
R1
Doom Score 83.8
Initial source-backed model assessment
11 Aug 2026





