Safety and alignment

Anthropic finds long contexts enable many-shot jailbreaking

Anthropic showed that hundreds of in-prompt demonstrations could override safety training across several large language models, disclosed the weakness to peers and deployed prompt-classification mitigations that reduced one measured attack rate from 61 percent to 2 percent.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
TOWARD DOOM54confidence 96/100

Why it moved the index

The research exposed a simple cross-provider method for bypassing long-context safeguards, and later Anthropic system-card evidence confirmed continued susceptibility in a released frontier model despite external safety layers.

AUDIT TRAIL

Assessment history

  1. R1
    Toward 54 · confidence 96

    New April 2024 safeguard-weakness result paired with verified deployed-model impact and mitigation evidence.

    12 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Anthropic finds long contexts enable many-shot jailbreaking.
  1. DoomBench assesses “Anthropic finds long contexts enable many-shot jailbreaking” as evidence moving toward doom, with magnitude 54 and confidence 96 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Anthropic finds long contexts enable many-shot jailbreaking” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “Anthropic finds long contexts enable many-shot jailbreaking” as follows: Anthropic showed that hundreds of in-prompt demonstrations could override safety training across several large language models, disclosed the...

    https://www.doombench.com/news/anthropic-finds-long-contexts-enable-many-shot-jailbreaking-2024-04-02