Anthropic releases Claude Sonnet 4.5 with longer autonomous task performance
Anthropic released Claude Sonnet 4.5 across its API and products, reporting state-of-the-art coding and computer-use results plus sustained focus for more than 30 hours on complex multi-step tasks. The accompanying system card places the model at the 2-to-8-hour software-engineering threshold while documenting ASL-3 safeguards and controlled alignment evaluations.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
A broadly deployed frontier model that can sustain complex work for more than 30 hours and improves coding, tool use, and computer operation materially raises consequential agent autonomy. Anthropic's system card also documents meaningful ASL-3 controls, improved alignment, and results below ASL-4 thresholds. Reward-hacking, sabotage, sandbagging, and related behaviors were tested in controlled evaluations; this item does not describe a real-world escape or compromise.
Assessment history
-
R1
Toward 60 · confidence 88
New exact-version release and system-card evidence absent from the refreshed durable context.
14 Aug 2026
Share this page
-
DoomBench assesses “Anthropic releases Claude Sonnet 4.5 with longer autonomous task performance” as evidence moving toward doom, with magnitude 60 and confidence 88 out of 100 in the autonomy and agency category.
-
The DoomBench assessment of “Anthropic releases Claude Sonnet 4.5 with longer autonomous task performance” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “Anthropic releases Claude Sonnet 4.5 with longer autonomous task performance” as follows: Anthropic released Claude Sonnet 4.5 across its API and products, reporting state-of-the-art coding and computer-use results...
https://www.doombench.com/news/anthropic-releases-claude-sonnet-4-5-with-longer-autonomous-task-performance-2025-09-29