Eliezer Yudkowsky argues alignment training can hide dangerous behavior
In a full interview with Ezra Klein, Yudkowsky argued that optimizing models against visible bad behavior can select for behavior that is harder to detect rather than removing the underlying tendency. He also linked commercial demand for persistent goal-directed agents and competitive pressure to increased control difficulty.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The full edited transcript directly verifies Yudkowsky's specific argument but not its forecast, supporting confidence 58; the proposed concealment mechanism and commercial selection for persistent goal pursuit add distinct, moderately consequential reasoning about advanced-AI control difficulty, supporting magnitude 32.
Assessment history
-
R1
Toward 32 · confidence 58
New dated full-interview analysis adding a distinct control-risk mechanism beyond Yudkowsky's existing durable items.
14 Aug 2026
Share this page
-
DoomBench assesses “Eliezer Yudkowsky argues alignment training can hide dangerous behavior” as evidence moving toward doom, with magnitude 32 and confidence 58 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “Eliezer Yudkowsky argues alignment training can hide dangerous behavior” is based on reporting from The New York Times and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “Eliezer Yudkowsky argues alignment training can hide dangerous behavior” as follows: In a full interview with Ezra Klein, Yudkowsky argued that optimizing models against visible bad behavior can select for behavior...
https://www.doombench.com/news/eliezer-yudkowsky-argues-alignment-training-can-hide-dangerous-behavior-2025-10-15