Anthropic traces planning, hidden goals and jailbreak circuits in Claude 3.5 Haiku
Anthropic's circuit-tracing work found forward planning, multilingual abstractions, fabricated reasoning, jailbreak dynamics and a controlled hidden-goal mechanism inside Claude 3.5 Haiku. The team later released the tracing tools for open-weight models and an interactive public interface.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The research exposed alignment-relevant mechanisms in a production model, tested causal interventions and was followed by a dated release of open-source circuit-tracing tooling. The authors state that useful explanations covered only about one quarter of attempted prompts, limiting assurance.
Assessment history
-
R1
Away 42 · confidence 92
Adds a missing Chris Olah-authored production-model interpretability result whose practical impact is verified by the subsequent public tool release.
14 Aug 2026
Share this page
-
DoomBench assesses “Anthropic traces planning, hidden goals and jailbreak circuits in Claude 3.5 Haiku” as evidence moving away from doom, with magnitude 42 and confidence 92 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “Anthropic traces planning, hidden goals and jailbreak circuits in Claude 3.5 Haiku” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “Anthropic traces planning, hidden goals and jailbreak circuits in Claude 3.5 Haiku” as follows: Anthropic's circuit-tracing work found forward planning, multilingual abstractions, fabricated reasoning, jailbreak...
https://www.doombench.com/news/anthropic-traces-planning-hidden-goals-and-jailbreak-circuits-in-claude-3-5-haiku-2025-03-27