Anthropic finds long contexts enable many-shot jailbreaking
Anthropic showed that hundreds of in-prompt demonstrations could override safety training across several large language models, disclosed the weakness to peers and deployed prompt-classification mitigations that reduced one measured attack rate from 61 percent to 2 percent.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The research exposed a simple cross-provider method for bypassing long-context safeguards, and later Anthropic system-card evidence confirmed continued susceptibility in a released frontier model despite external safety layers.
Assessment history
-
R1
Toward 54 · confidence 96
New April 2024 safeguard-weakness result paired with verified deployed-model impact and mitigation evidence.
12 Aug 2026
Share this page
-
DoomBench assesses “Anthropic finds long contexts enable many-shot jailbreaking” as evidence moving toward doom, with magnitude 54 and confidence 96 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “Anthropic finds long contexts enable many-shot jailbreaking” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “Anthropic finds long contexts enable many-shot jailbreaking” as follows: Anthropic showed that hundreds of in-prompt demonstrations could override safety training across several large language models, disclosed the...
https://www.doombench.com/news/anthropic-finds-long-contexts-enable-many-shot-jailbreaking-2024-04-02