Safety and alignment

US and UK tests find Claude safeguards routinely bypassed

The US and UK AI Safety Institutes reported that safeguards on the upgraded Claude 3.5 Sonnet could be circumvented in most US jailbreak tests and routinely circumvented in UK testing.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
TOWARD DOOM56confidence 99/100

Why it moved the index

Independent government predeployment testing identified direct control weaknesses in a frontier checkpoint across malicious-request safeguards, while the publication and cross-institute evaluation also strengthened external oversight.

AUDIT TRAIL

Assessment history

  1. R1
    Toward 56 · confidence 99

    New November 2024 independent evaluation of an existing exact checkpoint; no model revision is proposed.

    12 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for US and UK tests find Claude safeguards routinely bypassed.
  1. DoomBench assesses “US and UK tests find Claude safeguards routinely bypassed” as evidence moving toward doom, with magnitude 56 and confidence 99 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “US and UK tests find Claude safeguards routinely bypassed” is based on reporting from National Institute of Standards and Technology and records the editorial rationale, source quality, attribution, and...

  3. DoomBench summarizes “US and UK tests find Claude safeguards routinely bypassed” as follows: The US and UK AI Safety Institutes reported that safeguards on the upgraded Claude 3.5 Sonnet could be circumvented in most US jailbreak tests...

    https://www.doombench.com/news/us-and-uk-tests-find-claude-safeguards-routinely-bypassed-2024-11-19