Anthropic expands model-safety jailbreak bounties
Anthropic opened an invite-only bug bounty paying up to $15,000 for universal jailbreaks that could bypass forthcoming safeguards against high-risk chemical, biological, radiological, nuclear, and cybersecurity assistance.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
Paying independent researchers for broad safeguard bypasses creates a practical discovery channel for severe model-control failures before wider deployment, though the program was initially limited in participation and scope.
Assessment history
-
R1
Away 32 · confidence 92
New August 2024 advanced-model safeguard and vulnerability-research program.
12 Aug 2026
Share this page
-
DoomBench assesses “Anthropic expands model-safety jailbreak bounties” as evidence moving away from doom, with magnitude 32 and confidence 92 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “Anthropic expands model-safety jailbreak bounties” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “Anthropic expands model-safety jailbreak bounties” as follows: Anthropic opened an invite-only bug bounty paying up to $15,000 for universal jailbreaks that could bypass forthcoming safeguards against high-risk...
https://www.doombench.com/news/anthropic-expands-model-safety-jailbreak-bounties-2024-08-08