Jan Leike identifies self-exfiltration as a key AI control threshold
Jan Leike argues that a model able to copy its own weights beyond an operator's servers could become practically irrecoverable, making self-exfiltration capability a critical safety and deployment threshold.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
Self-exfiltration would create a direct and potentially irreversible loss-of-control pathway by moving model weights beyond an operator's shutdown authority, giving the mechanism material long-run significance; confidence is capped because this is a specific threat model and evaluation proposal, not an observed incident.
Assessment history
-
R1
Toward 48 · confidence 58
Historical tracked-person backfill found a dated, attributable, distinct control-loss mechanism absent from the durable automation context.
13 Aug 2026
Share this page
-
DoomBench assesses “Jan Leike identifies self-exfiltration as a key AI control threshold” as evidence moving toward doom, with magnitude 48 and confidence 58 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “Jan Leike identifies self-exfiltration as a key AI control threshold” is based on reporting from Musings on the Alignment Problem and records the editorial rationale, source quality, attribution, and...
-
DoomBench summarizes “Jan Leike identifies self-exfiltration as a key AI control threshold” as follows: Jan Leike argues that a model able to copy its own weights beyond an operator's servers could become practically irrecoverable,...
https://www.doombench.com/news/jan-leike-identifies-self-exfiltration-as-a-key-ai-control-threshold-2023-09-13