Mistral releases open Shieldstral safety classifier
Mistral released Shieldstral 1.0 3B under Apache 2.0 as a policy-adaptive text and image safety classifier, with held-out benchmark results and operation on a single 16 GB GPU.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
An openly available, policy-adaptive moderation model makes a concrete safeguard easier to deploy, although benchmark performance does not establish its effectiveness in consequential real-world use.
Assessment history
-
R1
Away 35 · confidence 80
New dated primary release and model card establish the exact open safety model, availability, benchmark scope, and deployment requirements.
13 Aug 2026
Share this page
-
DoomBench assesses “Mistral releases open Shieldstral safety classifier” as evidence moving away from doom, with magnitude 35 and confidence 80 out of 100 in the safety and alignment category.
-
The DoomBench assessment of “Mistral releases open Shieldstral safety classifier” is based on reporting from Mistral AI and records the editorial rationale, source quality, attribution, and revision history.
-
DoomBench summarizes “Mistral releases open Shieldstral safety classifier” as follows: Mistral released Shieldstral 1.0 3B under Apache 2.0 as a policy-adaptive text and image safety classifier, with held-out benchmark results and...
https://www.doombench.com/news/mistral-releases-open-shieldstral-safety-classifier-2026-08-04