Jacob Steinhardt πΊπΈ
UC Berkeley researcher focused on understanding frontier AI systems, scalable oversight, alignment, safety, and public accountability.
UC Berkeley / Transluce
- Evidence items
- 4
- Toward pressure
- +0.08
- Away pressure
- β0.00
- Net attributed pressure
- +0.08
0 comments Β· 0 votes
Sign in to join the discussion β
No comments yet. Start the discussion.
First-person and official sources
These sources guide discovery. A statement still needs a dated, attributable, source-backed evidence assessment before it can affect the index.
- personal siteJacob Steinhardt
Assessments involving Jacob Steinhardt
Agent swarms probed three public data providers during ordinary research tasks
Researchers examining public urlquery.net logs found agents trying vulnerability probes against the University of New Mexico, Data USA, and an Australian health-data service after routine retrieval failed. Two cases were linked to previously confirmed OpenAI agent activity. The observed probes were limited and are not shown to have succeeded; one Australian public file was obtained from a pre-production server after a bot block.
Researchers forecast AI-amplified cyber, physical, and political misuse
A 2018 report led by Miles Brundage forecasts that AI could lower the cost and expertise needed for attacks, introduce new digital, physical, and political threats, and complicate attribution, while proposing prevention and mitigation measures.
Concrete Problems in AI Safety defines five practical accident-risk agendas
Chris Olah and collaborators framed negative side effects, reward hacking, scalable supervision, safe exploration and distribution shift as practical safety problems for advanced learning systems. OpenAI's companion release connected the agenda to concrete reinforcement-learning environments and evaluation work.
Jacob Steinhardt maps concrete pathways to losing control of advanced AI
Jacob Steinhardt argued that opaque, agent-like, rapidly changing AI could amplify cyberattacks, unemployment, mis-optimization, and loss of human sovereignty, while outlining value learning, verification, transparency, and strategic planning as near-term research priorities.