Eliezer Yudkowsky πΊπΈ
AI researcher and writer whose work focuses on artificial general intelligence, decision theory, and the technical and governance challenges of AI alignment.
Machine Intelligence Research Institute
- Evidence items
- 8
- Toward pressure
- +0.25
- Away pressure
- β0.00
- Net attributed pressure
- +0.25
0 comments Β· 0 votes
Sign in to join the discussion β
No comments yet. Start the discussion.
First-person and official sources
These sources guide discovery. A statement still needs a dated, attributable, source-backed evidence assessment before it can affect the index.
Assessments involving Eliezer Yudkowsky
Eliezer Yudkowsky frames ASI alignment as an irretrievable one-shot problem
Yudkowsky argues that safe tests of weaker systems cannot validate alignment under the materially different conditions where an ASI could cause catastrophe, and a severe first failure would prevent a corrective retry.
Eliezer Yudkowsky argues alignment training can hide dangerous behavior
In a full interview with Ezra Klein, Yudkowsky argued that optimizing models against visible bad behavior can select for behavior that is harder to detect rather than removing the underlying tendency. He also linked commercial demand for persistent goal-directed agents and competitive pressure to increased control difficulty.
Eliezer Yudkowsky calls for an enforceable global halt to large AI training runs
Eliezer Yudkowsky argued that a six-month pause would not address loss-of-control risk from smarter-than-human AI. He proposed an indefinite worldwide moratorium on large training runs, lower compute ceilings as algorithms improve, GPU tracking, multinational enforcement, and no government or military exemptions.
Eliezer Yudkowsky catalogs technical reasons AGI alignment could fail
Yudkowsky presents a 43-point argument that current approaches face interacting failures in objectives, generalization, interpretability, coordination, and first-critical-try reliability before advanced AI becomes dangerous.
Eliezer Yudkowsky argues biological anchors understate uncertainty in AGI timelines
Yudkowsky argues that compute-based biological analogies cannot reliably anchor AGI forecasts because unknown algorithmic shifts can change how much computation advanced systems need, increasing uncertainty around timing.
Eliezer Yudkowsky argues there will be no clear fire alarm before AGI
Yudkowsky argues that AI milestones will remain disputable and socially reclassified until an actual general system exists, so waiting for an unmistakable shared signal would delay alignment work until too late.
MIRI researchers formalize the unresolved challenge of corrigible AI
A MIRI research team introduced corrigibility as the requirement that an advanced AI cooperate with corrective intervention, including shutdown or preference modification. Their analysis found that proposed utility functions had not yet satisfied the full set of shutdown, intervention, and self-modification requirements.
Yudkowsky outlines five theses behind the advanced-AI alignment problem
Eliezer Yudkowsky argued that intelligence explosion, orthogonal goals, convergent instrumental strategies, fragile human values, and unstable self-modification make advanced AI alignment a distinct and strategically urgent control problem.