Qwen2-VL-7B-Instruct
An open 7-billion-parameter visual-language model supporting images, multi-image inputs, long video, document understanding, and visual-agent operation of devices and robots.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why this model scores 57.3
Competitive multimodal reasoning, explicit device-operation capability, and downloadable weights combine strong practical autonomy with very broad deployment and limited downstream control.
News tied to Qwen2-VL-7B-Instruct
The model score of 57.3 rates this model's risk profile. The overall Doom Index of 67.9 measures the complete temporally weighted evidence record. These values answer different questions.
Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.
Alibaba releases Qwen2-VL visual-agent models
Alibaba released open 2B and 7B Qwen2-VL models and a 72B API model with image, long-video, multilingual text, and visual-agent capabilities for operating mobile devices and robots from visual inputs.
- Full item contribution
- +0.12
- Qwen2-VL-7B-Instruct equal share
- +0.04
Model score history
-
R3
Doom Score 57.3
Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-142416 under temporal-monthly-pressure-v4.
12 Aug 2026 -
R2
Doom Score 57.8
Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-132112 under fixed-sensitivity-v3.
12 Aug 2026 -
R1
Doom Score 58.9
New exact open Qwen2-VL model tier verified by the dated primary launch.
12 Aug 2026