Safety and alignment

Mistral releases open Shieldstral safety classifier

Mistral released Shieldstral 1.0 3B under Apache 2.0 as a policy-adaptive text and image safety classifier, with held-out benchmark results and operation on a single 16 GB GPU.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM35confidence 80/100

Why it moved the index

An openly available, policy-adaptive moderation model makes a concrete safeguard easier to deploy, although benchmark performance does not establish its effectiveness in consequential real-world use.

AUDIT TRAIL

Assessment history

  1. R1
    Away 35 · confidence 80

    New dated primary release and model card establish the exact open safety model, availability, benchmark scope, and deployment requirements.

    13 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Mistral releases open Shieldstral safety classifier.
  1. DoomBench assesses “Mistral releases open Shieldstral safety classifier” as evidence moving away from doom, with magnitude 35 and confidence 80 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Mistral releases open Shieldstral safety classifier” is based on reporting from Mistral AI and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “Mistral releases open Shieldstral safety classifier” as follows: Mistral released Shieldstral 1.0 3B under Apache 2.0 as a policy-adaptive text and image safety classifier, with held-out benchmark results and...

    https://www.doombench.com/news/mistral-releases-open-shieldstral-safety-classifier-2026-08-04