Victoria Krakovna π¨π¦ π¬π§
AI-safety researcher whose work covers specification gaming, reward misspecification, agent incentives, and practical alignment methods.
- Evidence items
- 4
- Toward pressure
- +0.03
- Away pressure
- β0.00
- Net attributed pressure
- +0.03
0 comments Β· 0 votes
Sign in to join the discussion β
No comments yet. Start the discussion.
Current and former organisations
These dated, source-backed roles support navigation between people and companies. They do not attribute a story or change the Doom Index.
First-person and official sources
These sources guide discovery. A statement still needs a dated, attributable, source-backed evidence assessment before it can affect the index.
- personal siteVictoria Krakovna
Assessments involving Victoria Krakovna
Victoria Krakovna reframes AI alignment as a near-term safety problem
Victoria Krakovna argued that advanced-AI alignment should be treated as medium-term safety affecting people alive today, not only as a longtermist concern. She identified pandemics and economic takeover as plausible catastrophic pathways and called for safety evaluations, governance standards, and slow, cautious deployment.
Victoria Krakovna maps AI specification failures to four Goodhart effects
Victoria Krakovna and Ramana Kumar map objective-specification failures to regressional, extremal, causal, and adversarial Goodhart effects. Their framework distinguishes gaps between ideal, model, design, implementation, and revealed objectives, linking them to specification gaming, side effects, reward tampering, robustness failures, and deceptive alignment.
Victoria Krakovna tests relative reachability as a safeguard against harmful side effects
Victoria Krakovna reports that a relative-reachability penalty avoided two incentive failures in tabular AI Safety Gridworlds: blocking irreversible events and undoing helpful interventions. The proof of concept preserved safe options better than reversibility and simple impact penalties, while leaving realistic scale and default-outcome definitions unresolved.
Victoria Krakovna argues AGI risk does not require an intelligence explosion
Victoria Krakovna argues that general AI could create existential risk even without rapid recursive self-improvement. She identifies competitive development incentives, convergent instrumental goals, unintended objectives, difficult value learning, containment failure, and slow institutional coordination as independent control-risk pathways.