Beth Barnes πΊπΈ
Founder and CEO of METR, where she leads independent evaluations of frontier-model autonomy, misalignment, and catastrophic risk. She previously conducted alignment and safety research at OpenAI and Google DeepMind.
METR
- Evidence items
- 7
- Toward pressure
- +0.12
- Away pressure
- β0.36
- Net attributed pressure
- -0.24
0 comments Β· 0 votes
Sign in to join the discussion β
No comments yet. Start the discussion.
Current and former organisations
These dated, source-backed roles support navigation between people and companies. They do not attribute a story or change the Doom Index.
First-person and official sources
These sources guide discovery. A statement still needs a dated, attributable, source-backed evidence assessment before it can affect the index.
- interview feed80,000 Hours interview with Beth Barnes
- institutional profileBeth Barnes at METR
- publication feedBeth Barnes on LessWrong
- personal siteBeth Barnes personal site
Assessments involving Beth Barnes
METR finds low sabotage risk in Claude Opus 4 and 4.1 after independent review
METR reviewed unredacted evidence and agreed that catastrophic sabotage risk from Claude Opus 4 and 4.1 was low, while warning that reasoning-hiding and evaluation-awareness claims were overconfident. Anthropic said the review led it to revise threat models and alignment practices, including training changes for reasoning monitorability.
Beth Barnes says AI labs are locally reasonable but globally reckless
In a full interview, Barnes argued that competitive incentives can make individually understandable lab decisions collectively unsafe. She said large uncertainty around threat models, capability measurement, and meaningful external oversight makes even a one-percent catastrophic-risk target difficult to justify.
ARC Evals becomes METR and begins separating into an independent nonprofit
ARC Evals adopted the METR name while wrapping up its incubation at the Alignment Research Center and transitioning toward a standalone nonprofit focused on measuring frontier-model autonomy and threat-relevant capabilities under Beth Barnes's leadership.
ARC Evals finds 2023 agents far below autonomous replication threshold
ARC Evals tested four GPT-4- and Claude-based agents on 12 realistic autonomous replication and adaptation tasks. The agents completed only the easiest tasks and failed end-to-end persistent deployment and controlled phishing attempts, indicating that casual users of those versions were unlikely to create dangerous autonomous agents. OpenAI's GPT-4 system card independently documents ARC's predeployment evaluation role.
Beth Barnes builds an ARC team to evaluate model power-seeking capabilities
Beth Barnes described a new Alignment Research Center team already building capability evaluations for advanced language models. The work targeted long-horizon agency, resource acquisition, oversight evasion, and thresholds that labs could use to pause scaling or deployment until stronger alignment measures existed.
Beth Barnes outlines how persuasive AI could amplify manipulation before AGI
Beth Barnes argued that personalized AI companions and assistants could make persuasion easier to optimize at scale, potentially strengthening ideological lock-in, authoritarian control, and poor collective decision-making. She treated the pathway as uncertain and smaller than standard alignment failure, but directly linked it to deception and loss of human control.
Barnes identifies obfuscated arguments as a scalable-oversight failure mode
Beth Barnes reports human debate experiments in which dishonest debaters could construct large arguments containing rare fatal errors that honest debaters and judges could not reliably locate. She says the team had no fix and that the result may place an important quantitative limit on debate and iterated amplification as methods for supervising decisions beyond unaided human verification.