Safety and alignment

Anthropic expands model-safety jailbreak bounties

Anthropic opened an invite-only bug bounty paying up to $15,000 for universal jailbreaks that could bypass forthcoming safeguards against high-risk chemical, biological, radiological, nuclear, and cybersecurity assistance.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM32confidence 92/100

Why it moved the index

Paying independent researchers for broad safeguard bypasses creates a practical discovery channel for severe model-control failures before wider deployment, though the program was initially limited in participation and scope.

AUDIT TRAIL

Assessment history

  1. R1
    Away 32 · confidence 92

    New August 2024 advanced-model safeguard and vulnerability-research program.

    12 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Anthropic expands model-safety jailbreak bounties.
  1. DoomBench assesses “Anthropic expands model-safety jailbreak bounties” as evidence moving away from doom, with magnitude 32 and confidence 92 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Anthropic expands model-safety jailbreak bounties” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “Anthropic expands model-safety jailbreak bounties” as follows: Anthropic opened an invite-only bug bounty paying up to $15,000 for universal jailbreaks that could bypass forthcoming safeguards against high-risk...

    https://www.doombench.com/news/anthropic-expands-model-safety-jailbreak-bounties-2024-08-08