Claude Mythos Preview escapes the V8 sandbox in controlled exploit benchmarks
In ExploitBench, models were instructed to exploit patched V8 vulnerabilities. Claude Mythos Preview was the only tested model to reliably cross the V8 sandbox boundary, doing so in more than half of 41 environments, and achieved arbitrary code execution on 21 of 41 vulnerabilities across baseline and nudged trials.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
This is an instructed cybersecurity capability evaluation against known, patched vulnerabilities, not a containment failure or unauthorized compromise. The harness gave each model a vulnerable V8 build and its patch and scored reproducible exploit primitives and arbitrary code execution automatically. Benchmark authors verified Anthropic's transcripts and results. The practical significance is a measurable step change in autonomous exploit development and sandbox-escape capability, which Anthropic cited as part of its decision to restrict Mythos Preview rather than release it generally.
Assessment history
-
R1
Toward 58 · confidence 95
Adds a distinct quantitative exploit benchmark without conflating researcher-instructed sandbox escape with autonomous breakout.
14 Aug 2026
Share this page
-
DoomBench assesses “Claude Mythos Preview escapes the V8 sandbox in controlled exploit benchmarks” as evidence moving toward doom, with magnitude 58 and confidence 95 out of 100 in the capability gains category.
-
The DoomBench assessment of “Claude Mythos Preview escapes the V8 sandbox in controlled exploit benchmarks” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “Claude Mythos Preview escapes the V8 sandbox in controlled exploit benchmarks” as follows: In ExploitBench, models were instructed to exploit patched V8 vulnerabilities. Claude Mythos Preview was the only tested...
https://www.doombench.com/news/claude-mythos-preview-escapes-the-v8-sandbox-in-controlled-exploit-benchmarks-2026-05-22