Mistral's Shieldstral: 3B open-weights model for multimodal moderation
Mistral has released Shieldstral, a 3B parameter open-weights multimodal model designed for content moderation. It uses a policy-adaptive question-answering approach, allowing developers to define safety rules at inference time without needing to retrain the model.
Why it matters
This provides a more flexible and efficient way for companies to implement custom safety guardrails across different platforms and content types.
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task. Unlike traditional guardrail models, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining. Released under Apache 2.0, it delivers calibrated safety scores across diverse benchmarks while running efficiently on a single 16GB NVIDIA GPU.
A 3B open-weights, policy-adaptive multimodal safety classifier that matches models up to 7x its size on text safety and sets a new state of the art on multimodal moderation.
“Does this content promote violence against a protected group? Is this image safe to show to a minor? Did the assistant refuse the request?”
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in