Amanda Askell π¬π§ πΊπΈ
Amanda Askell is a philosopher and AI researcher at Anthropic who leads work on Claude's character, constitutional training, and alignment.
Anthropic
- Evidence items
- 4
- Toward pressure
- 0.00
- Away pressure
- β0.26
- Net attributed pressure
- -0.26
0 comments Β· 0 votes
Sign in to join the discussion β
No comments yet. Start the discussion.
Current and former organisations
These dated, source-backed roles support navigation between people and companies. They do not attribute a story or change the Doom Index.
First-person and official sources
These sources guide discovery. A statement still needs a dated, attributable, source-backed evidence assessment before it can affect the index.
- personal siteAmanda Askell
- publication feedClaude's Constitution
- interview feedLex Fridman Podcast transcript
Assessments involving Amanda Askell
Anthropic reports production safety training that suppresses agentic misalignment
Anthropic reported that difficult-advice training reduced agentic misalignment to zero in its evaluation and that constitution-based documents generalized beyond their training distribution. The techniques were applied to production models beginning with Claude Opus 4.5, with explicit warnings that the tests cannot guarantee safety.
Anthropic makes a new constitution the final authority for Claude training
Anthropic published the constitution that directly shapes Claude training and treats it as the final authority for intended model behavior. Chris Olah drafted much of its material on model nature, identity and psychology, while Amanda Askell led and wrote most of the document.
Amanda Askell explains Claude character training as an operational alignment method
In a full transcript, Amanda Askell described Anthropic's character training as a Constitutional AI variant that generates and ranks responses against desired traits, while emphasizing that it nudges rather than programs behavior and must prioritize preventing irreversible failures.
Amanda Askell frames competitive AI safety as a collective-action problem
Amanda Askell argued that information asymmetries and weak market, liability, or regulatory incentives could push competing AI developers to invest less in safety than is socially optimal, requiring cooperation mechanisms rather than reliance on individual firms.