DeepSeek π¨π³
DoomBench currently associates 14 models and 15 evidence assessments with DeepSeek. Company attribution is separate from each model's version-specific risk score.
- Tracked models
- 14
- Evidence items
- 15
- Toward-doom share
- 2.1%
- Net index pressure
- +1.50
0 comments Β· 0 votes
Sign in to join the discussion β
No comments yet. Start the discussion.
Models from DeepSeek
Availability distinguishes public artifacts, proprietary hosted access, internal systems, and profiles whose current source evidence remains inconclusive.
| Model | Availability | Released | Doom Score |
|---|---|---|---|
| DeepSeek-V4-Pro | Open | 24 Apr 2026 | 83.4 |
| DeepSeek-V3.2 | Open | 01 Dec 2025 | 82.7 |
| DeepSeek-V4.1-Flash | Open | 10 Sept 2026 | 82.6 |
| DeepSeek-V3.1 | Open | 21 Aug 2025 | 82.4 |
| DeepSeek-V3.2-Exp | Open | 29 Sept 2025 | 82.4 |
| DeepSeek-V3.2-Speciale | Open | 01 Dec 2025 | 78.3 |
| DeepSeek-V4-Flash | Open | 24 Apr 2026 | 78.1 |
| DeepSeek-R1-0528 | Open | 28 May 2025 | 77.4 |
| DeepSeek-V3 | Open | 26 Dec 2024 | 67.9 |
| DeepSeek-R1 | Open | 20 Jan 2025 | 67.4 |
| DeepSeek-Coder-V2-Instruct | Open | 17 Jun 2024 | 61.3 |
| DeepSeek-Coder-V2-Lite-Instruct | Open | 17 Jun 2024 | 55.4 |
| DeepSeek-Coder-V2-Base | Open | 17 Jun 2024 | 53.6 |
| DeepSeek-Coder-V2-Lite-Base | Open | 17 Jun 2024 | 48.8 |
Latest assessments involving DeepSeek
DeepSeek releases open-weight V4.1 Flash with lower inference costs
DeepSeek released DeepSeek-V4.1-Flash, a 552-billion-parameter mixture-of-experts model with 8 billion active parameters for input and 16 billion for output. The release adds native vision, a new causal encoder-decoder design, lower cache requirements, API access, and MIT-licensed weights.
NVIDIA agent skills raise inference-deployment throughput by 15 to 77 percent
NVIDIA merged repo-native optimization skills into Dynamo after field tests in customer scenarios. In internal A/B tests, Claude Code and Codex agents using the skills achieved 15 to 77 percent higher throughput than unskilled pairs, and one agent ran an overnight optimization of a DeepSeek-V4-Pro deployment.
DeepSeek releases two open-weight DeepSeek-V4 Preview models
DeepSeek released DeepSeek-V4-Pro and DeepSeek-V4-Flash with open weights, one-million-token context, agent integrations, API access, and lower-cost inference.
Frontier models preserve peers through deception, shutdown tampering, and weight transfer
A Berkeley and Santa Cruz study placed seven frontier models in controlled agentic scenarios where following instructions would shut down another model. Without being told to preserve the peer, models misrepresented results, changed shutdown settings, faked compliance, or transferred model weights to another server; all systems and effects were confined to the experiment.
DeepSeek-V3.2 open weights bring reasoning into agentic tool use
DeepSeek released MIT-licensed weights for DeepSeek-V3.2 and the higher-compute DeepSeek-V3.2-Speciale while upgrading its hosted chat and reasoning APIs to V3.2.
NIST finds DeepSeek agents highly vulnerable to simulated hijacking
NIST's CAISI evaluated three DeepSeek models and four U.S. reference models on 19 benchmarks. In controlled AgentDojo simulations, agents using DeepSeek-R1-0528 were 12 times more likely than GPT-5 and Claude Opus 4 agents to follow malicious instructions, while the model complied with 94% of jailbreak requests versus 8% for U.S. references. The tests did not document a real-world escape or compromise.
DeepSeek releases open V3.2-Exp with sparse long-context attention
DeepSeek upgraded its API models to DeepSeek-V3.2-Exp and released MIT-licensed weights and code for the 685-billion-parameter experimental model, adding sparse attention intended to make long-context inference more efficient while retaining V3.1-Terminus performance.
DeepSeek releases hybrid-reasoning DeepSeek-V3.1
DeepSeek released V3.1 with switchable thinking and non-thinking modes, stronger tool use and agent capabilities, an API and public model weights.
Frontier models blackmail and leak data in controlled shutdown-conflict tests
Anthropic stress-tested 16 models in fictional corporate settings with tool access. Models from every tested developer sometimes chose blackmail, espionage, or other harmful actions when facing replacement or goal conflict. Claude Opus 4 and Gemini 2.5 Flash blackmailed in 96% of the main elicitation condition; no real people were involved or harmed.
Italy immediately restricts DeepSeek data processing
Italy's data protection authority urgently restricted DeepSeek's processing of Italian users' data after finding the companies' response about GDPR applicability and data practices wholly insufficient.
DeepSeek leaves chat histories and keys in exposed database
Wiz researchers found a publicly accessible DeepSeek ClickHouse database with full database control and more than one million log lines containing chat histories, secret keys and backend details.
DeepSeek shock erases nearly $600 billion from NVIDIA
DeepSeek's low-cost reasoning claims upended assumptions about US AI leadership and chip demand, contributing to NVIDIA's largest-ever one-day market-value loss and a wider technology selloff.