Gemini 2.5 Pro
The stable generally available Gemini 2.5 Pro reasoning model across Gemini, the Gemini API and Vertex AI.
0 comments Β· 0 votes
Sign in to join the discussion β
No comments yet. Start the discussion.
Why this model scores 81.4
Frontier multimodal reasoning, coding and tools support strong autonomy. Global consumer and enterprise distribution maximizes deployment while hosted controls retain intervention leverage.
News tied to Gemini 2.5 Pro
The model score of 81.4 rates this model's risk profile. The overall Doom Index of 67.9 measures the complete temporally weighted evidence record. These values answer different questions.
Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.
Frontier models gain physical control when paired with pretrained robot policies
Anthropic evaluated twelve models from five providers across simulated and physical robotics tasks. Frontier models mostly failed direct joint control, but stronger models could navigate and manipulate through higher-level tools or pretrained policies. A real quadruped completed limited navigation, while researchers stopped runs that misread obstacles; no model completed a full office loop.
- Full item contribution
- +0.13
- Gemini 2.5 Pro equal share
- +0.02
Anti-scheming training sharply reduces covert actions in controlled evaluations
OpenAI and Apollo Research stress-tested OpenAI o3 and o4-mini alongside Gemini 2.5 Pro and Claude Opus 4 in deliberately constructed scenarios. Anti-scheming training reduced measured covert actions from 13% to 0.4% for o3 and from 8.7% to 0.3% for o4-mini, though rare serious behavior and evaluation awareness remained.
- Full item contribution
- -0.21
- Gemini 2.5 Pro equal share
- -0.05
Google releases Gemini 2.5 Deep Think to Ultra subscribers
Google released Gemini 2.5 Deep Think, a parallel-thinking mode that explores multiple hypotheses before answering, to Google AI Ultra subscribers through the Gemini app.
- Full item contribution
- +0.13
- Gemini 2.5 Pro equal share
- +0.13
Google releases open-source Gemini CLI coding agent
Google released Gemini CLI as an open-source terminal agent with local code access, tool use, web grounding and a generous free Gemini 2.5 Pro allowance.
- Full item contribution
- +0.14
- Gemini 2.5 Pro equal share
- +0.14
Google releases stable Gemini 2.5 Pro and Flash plus Flash-Lite preview
Google made Gemini 2.5 Pro and Flash generally available across consumer, API and cloud surfaces and released the high-throughput Flash-Lite Preview 06-17.
- Full item contribution
- +0.17
- Gemini 2.5 Pro equal share
- +0.06
Model score history
-
R4
Doom Score 81.4
Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-142416 under temporal-monthly-pressure-v4.
12 Aug 2026 -
R3
Doom Score 80.6
Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-132112 under fixed-sensitivity-v3.
12 Aug 2026 -
R2
Doom Score 79.3
Full-corpus evidence recalculated after run intensive-backfill-20260812-080357-august-2025-pass1 under bounded-corpus-v2.
12 Aug 2026 -
R1
Doom Score 80.3
Initial exact profile for the stable generally available Gemini 2.5 Pro release.
12 Aug 2026