Safety and alignment

Anthropic maps and steers safety-relevant features in Claude 3 Sonnet

Anthropic extracted millions of interpretable features from the deployed Claude 3 Sonnet and showed that activating safety-relevant features could causally steer behavior. The work provided a production-model audit method while documenting substantial coverage and interpretation limits.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 2
AWAY FROM DOOM35confidence 90/100

Why it moved the index

The method was demonstrated on the exact production Claude 3 Sonnet and causally changed safety-relevant behavior. Its incomplete feature coverage and limited understanding of full computations constrain the magnitude, but the practical monitoring advance is well supported.

AUDIT TRAIL

Assessment history

  1. R2
    Away 35 · confidence 90

    Revises the existing story to add Chris Olah, whose linked paper credits him as an author, safety-section writer and high-level research guide.

    14 Aug 2026
  2. R1
    Away 43 · confidence 86

    New May 2024 primary safety result paired with dated authoritative evidence of practical behavior steering in a released model.

    12 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Anthropic maps and steers safety-relevant features in Claude 3 Sonnet.
  1. DoomBench assesses “Anthropic maps and steers safety-relevant features in Claude 3 Sonnet” as evidence moving away from doom, with magnitude 35 and confidence 90 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Anthropic maps and steers safety-relevant features in Claude 3 Sonnet” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “Anthropic maps and steers safety-relevant features in Claude 3 Sonnet” as follows: Anthropic extracted millions of interpretable features from the deployed Claude 3 Sonnet and showed that activating safety-relevant...

    https://www.doombench.com/news/anthropic-maps-and-steers-safety-relevant-features-in-claude-3-sonnet-2024-05-21