Eliezer Yudkowsky frames ASI alignment as an irretrievable one-shot problem
Yudkowsky argues that safe tests of weaker systems cannot validate alignment under the materially different conditions where an ASI could cause catastrophe, and a severe first failure would prevent a corrective retry.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The post adds a distinct technical mechanism: distribution shift separates survivable development tests from the first capability regime in which alignment failure could remove the possibility of recovery. It is a detailed first-person argument, but not an empirical demonstration that the predicted critical transition will occur.
Assessment history
-
R1
Toward 22 · confidence 56
Initial inclusion from a dated first-person source recovered during Eliezer Yudkowsky's historical backfill.
13 Aug 2026
Share this page
-
DoomBench assesses “Eliezer Yudkowsky frames ASI alignment as an irretrievable one-shot problem” as evidence moving toward doom, with magnitude 22 and confidence 56 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “Eliezer Yudkowsky frames ASI alignment as an irretrievable one-shot problem” is based on reporting from LessWrong and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “Eliezer Yudkowsky frames ASI alignment as an irretrievable one-shot problem” as follows: Yudkowsky argues that safe tests of weaker systems cannot validate alignment under the materially different conditions where...
https://www.doombench.com/news/eliezer-yudkowsky-frames-asi-alignment-as-an-irretrievable-one-shot-problem-2026-05-04