Safety and alignment

Anthropic stress-tests ASL-3 jailbreak defenses before deployment

Anthropic launched an invite-only bug bounty to find universal jailbreaks in Constitutional Classifiers designed to block biological-weapons assistance under its ASL-3 deployment standard.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM37confidence 78/100

Why it moved the index

External adversarial testing of severe biological-risk classifiers strengthens deployment controls before frontier release, although the program was narrow, invite-only and provider-governed.

AUDIT TRAIL

Assessment history

  1. R1
    Away 37 · confidence 78

    New operational severe-risk safeguard test with no durable collision.

    12 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Anthropic stress-tests ASL-3 jailbreak defenses before deployment.
  1. DoomBench assesses “Anthropic stress-tests ASL-3 jailbreak defenses before deployment” as evidence moving away from doom, with magnitude 37 and confidence 78 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Anthropic stress-tests ASL-3 jailbreak defenses before deployment” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “Anthropic stress-tests ASL-3 jailbreak defenses before deployment” as follows: Anthropic launched an invite-only bug bounty to find universal jailbreaks in Constitutional Classifiers designed to block...

    https://www.doombench.com/news/anthropic-stress-tests-asl-3-jailbreak-defenses-before-deployment-2025-05-14