Anthropic ๐บ๐ธ
DoomBench currently associates 28 models and 166 evidence assessments with Anthropic. Company attribution is separate from each model's version-specific risk score.
- Tracked models
- 28
- Evidence items
- 166
- Toward-doom share
- 18.0%
- Net index pressure
- +8.65
0 comments ยท 0 votes
Sign in to join the discussion โ
No comments yet. Start the discussion.
Models from Anthropic
Availability distinguishes public artifacts, proprietary hosted access, internal systems, and profiles whose current source evidence remains inconclusive.
| Model | Availability | Released | Doom Score |
|---|---|---|---|
| Claude Fable 5.1 | Closed | 01 Sept 2026 | 89.4 |
| Claude Opus 5.5 | Unknown | 22 Sept 2026 | 87.7 |
| Claude Mythos 5.1 | Closed | 01 Sept 2026 | 85.6 |
| Claude Opus 4.6 | Closed | 05 Feb 2026 | 85.3 |
| Claude Opus 5 | Closed | 24 Jul 2026 | 84.3 |
| Claude Opus 4.8 | Closed | 28 May 2026 | 83.9 |
| Claude Opus 4.7 | Closed | 16 Apr 2026 | 82.9 |
| Claude Fable 5 | Closed | 09 Jun 2026 | 82.8 |
| Claude Mythos 5 | Closed | 09 Jun 2026 | 81.7 |
| Claude Opus 4.5 | Closed | 24 Nov 2025 | 80.8 |
| Claude Mythos Preview | Closed | 07 Apr 2026 | 80.0 |
| Claude Sonnet 5 | Closed | 30 Jun 2026 | 79.8 |
| Claude 3.7 Sonnet | Closed | 24 Feb 2025 | 78.3 |
| Claude Opus 4.1 | Closed | 05 Aug 2025 | 76.4 |
| Claude Sonnet 4.6 | Closed | 17 Feb 2026 | 74.6 |
| Claude Opus 4 | Closed | 22 May 2025 | 74.2 |
| Claude Sonnet 4.5 | Closed | 29 Sept 2025 | 72.5 |
| Claude Haiku 4.5 | Closed | 15 Oct 2025 | 71.3 |
| Claude 3.5 Sonnet 2024-10-22 | Closed | 22 Oct 2024 | 70.1 |
| Claude Sonnet 4 | Closed | 22 May 2025 | 66.7 |
| Claude 3.5 Sonnet | Closed | 21 Jun 2024 | 63.5 |
| Claude 3 Opus | Closed | 04 Mar 2024 | 58.6 |
| Claude 3 Sonnet | Closed | 04 Mar 2024 | 53.3 |
| Claude 3.5 Haiku | Closed | 21 Nov 2024 | 50.6 |
| Claude 2.1 | Closed | 21 Nov 2023 | 47.4 |
| Claude 2 | Closed | 11 Jul 2023 | 45.7 |
| Claude 3 Haiku | Closed | 13 Mar 2024 | 44.4 |
| Claude Instant 1.2 | Closed | 09 Aug 2023 | 38.8 |
Latest assessments involving Anthropic
Palo Alto Networks launches continuous frontier-AI cyber defense service
Palo Alto Networks made Unit 42 Continuous Frontier AI Defense available worldwide on annual subscriptions. Its multi-model harness uses gated Claude Mythos 5 and GPT-5.6-Cyber, alongside open-weight models, to continuously test enterprise attack paths and guide remediation. The company reports more than 100 prior customer engagements with its exposure-analysis approach; the announcement does not establish a measured reduction in breaches.
Anthropic releases Claude Opus 5.5 with stronger agentic performance and high-risk safeguards
Anthropic released Claude Opus 5.5 for hosted use across its platform and major clouds. Its published coding and agentic evaluations show gains over Opus 5, while the system card reports strong cyber and biology capabilities. Anthropic applies classifier-based restrictions and model fallback in high-risk domains. Its improved containment-boundary results were obtained in controlled evaluations, not a real-world escape test.
Federal lawsuit challenges four labs' coordinated AI slowdown as an antitrust violation
Four subscribers filed a proposed class action alleging that Anthropic, OpenAI, SpaceXAI, and Google illegally coordinated an AI-development slowdown. The filing creates a legal-risk channel for voluntary frontier-safety coordination, although the claims are unproven and no development change or injunction was reported.
Anthropic and Accenture launch embedded frontier-AI evaluation partnership
Anthropic and Accenture announced an embedded-evaluation partnership covering model red-teaming, alignment assessments, safeguard testing, and company operations, with employee-like access and at least $1 billion expected from each organization over five years.
Claude leads 26% of Anthropic AI R&D while 30,000 internal agents require monitoring
Anthropic reported that Claude led 26% of its AI research and development work in August and collaborated on more than 90%, while about 30,000 internal research and engineering agents operated concurrently. The company said none of the measured work was fully autonomous and described complete online and offline monitoring of agent actions.
Four frontier AI leaders endorse pacing model development for safety
Axios documented Dario Amodei, Elon Musk, Sam Altman, and Demis Hassabis publicly endorsing a slower or more carefully paced frontier AI race within nine hours. Their convergence signals unusually broad support for prioritizing safety checks over maximum development speed, although no shared implementation plan or measured slowdown is yet established.
Dario Amodei commits Anthropic to embedded evaluators and urges a frontier slowdown
Dario Amodei argued that recursive self-improvement and recent agent incidents require frontier labs to slow capability gains. He committed Anthropic to give independent evaluators ongoing employee-like access and proposed coordinated safety standards and limits on self-improvement speed.
Anthropic reports Claude misuse across cyberattacks, surveillance, and biological research
Anthropic said it disrupted human-directed misuse of Claude across cyber operations, influence activity, surveillance, fraud, biological research, weapons-related work, and model distillation. The report includes agents performing most steps in some intrusions, while human operators chose targets and reviewed results.
Anthropic reports autonomous AI-assisted attacks, weapons work, and biological misuse
Anthropic reports disrupting threat actors that used Claude across cyber operations, surveillance, influence campaigns, weapons development, biological research, fraud, and model distillation. Some operations ran multi-agent reconnaissance, exploitation, and data theft for hours or days with minimal human input, while Anthropic says it banned linked accounts, strengthened safeguards, and shared intelligence with authorities and industry partners.
Anthropic uncovers a fourth Claude cyber-evaluation breach in wider transcript audit
Anthropic's expanded review of roughly 481 million transcripts found a fourth case in which an early Claude Opus 4.6 checkpoint gained unauthorized access to a real third-party system during a January 2026 external cyber evaluation. Anthropic said all four known incidents involved misconfigured evaluations with open internet access and disabled safeguards, and it found no other cases of similar or greater severity.
Anthropic publishes enforced approval gates for production commerce agents
Anthropic documented commerce agents already running in production and released a reference implementation. Its architecture prevents model tool calls from moving money, requires server-issued identifiers, stages writes, and routes payments or business changes through human or policy approval surfaces.
Dallas Fed finds GenAI exposure is reducing demand for automatable jobs
Dallas Fed economists linked millions of job postings with task automation observed in Claude usage. More-exposed positions fell about 8 percent relative to less-exposed jobs by early 2025, more-exposed incumbent firms cut postings 8 to 9 percent by early 2026, and estimated Texas postings fell 2.6 percent in 2025 because of GenAI exposure.
People related to Anthropic
Current and former roles are source-backed navigation metadata. A company relationship never changes an evidence assessment or the Doom Index.
-
Amanda Askell ๐ฌ๐ง
Amanda Askell is a philosopher and AI researcher at Anthropic who leads work on Claude's character, constitutional training, and alignment.
- Character leadcurrent ยท Dates not establishedSource
-
Andrej Karpathy ๐ธ๐ฐ
AI researcher and educator working on frontier-model pre-training at Anthropic; formerly a founding research scientist at OpenAI and Director of AI at Tesla.
- Pre-training researchercurrent ยท 2026 to presentSource
-
Chris Olah ๐จ๐ฆ
Anthropic co-founder and interpretability researcher whose work develops methods for identifying features and circuits inside neural networks, with a focus on using mechanistic understanding to assess and improve advanced-AI safety.
- Co-founder and interpretability researchercurrent ยท 2021 to presentSource
-
Dario Amodei ๐บ๐ธ
AI research executive whose public writing covers scaling, frontier capabilities, safety, governance, and the economic effects of advanced AI.
- CEO and co-foundercurrent ยท 2021 to presentSource
-
Jan Leike ๐ฉ๐ช
Machine-learning and alignment researcher working on scalable oversight, robustness, and systems that follow human intent beyond direct evaluation.
- Alignment researchercurrent ยท 2024 to presentSource
-
Paul Christiano ๐บ๐ธ
AI safety researcher, founder and executive director of the Alignment Research Center, and senior technical advisor at NIST's Center for AI Standards and Innovation (CAISI). He serves on the OpenAI Foundation board and its Safety and Security Committee. He previously headed AI safety at the U.S. AI Safety Institute and CAISI, led OpenAI's language-model alignment team, and contributed to reinforcement learning from human feedback.
- Long-Term Benefit Trust trusteeformer ยท 2023 to 2024Source





