US and UK safety institutes expose and fix agent attack paths before deployment
OpenAI reported authorized US CAISI and UK AISI testing of GPT-5 and ChatGPT Agent. A controlled proof-of-concept exploit chain succeeded in roughly half of trials before OpenAI fixed the issue within one business day, while other findings changed product, policy, and training controls.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
This was controlled, permissioned red-team testing, not a real-world escape or external compromise. It is material because the evaluators demonstrated an end-to-end attack path against a deployed agent configuration and the work produced a verified operational response: rapid remediation plus broader product, policy, and training changes.
Assessment history
-
R1
Away 44 · confidence 92
New historical evidence found in the September 2025 gap review.
13 Aug 2026
Share this page
-
DoomBench assesses “US and UK safety institutes expose and fix agent attack paths before deployment” as evidence moving away from doom, with magnitude 44 and confidence 92 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “US and UK safety institutes expose and fix agent attack paths before deployment” is based on reporting from OpenAI and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “US and UK safety institutes expose and fix agent attack paths before deployment” as follows: OpenAI reported authorized US CAISI and UK AISI testing of GPT-5 and ChatGPT Agent. A controlled proof-of-concept exploit...
https://www.doombench.com/news/us-and-uk-safety-institutes-expose-and-fix-agent-attack-paths-before-deployment-2025-09-12