Mistral's Shieldstral: 3B open-weights model for multimodal moderation
Mistral released Shieldstral, a 3B open-weights model for text and image safety classification that operates as a question-answering system accepting plain-language policies without retraining. The model matches performance with models up to 7 times larger, runs on a single GPU, and is available under Apache 2.0 license.
How does Shieldstral differ from traditional content moderation models?
Shieldstral accepts moderation policies as natural language text at inference time, enabling adaptation to different contexts without retraining. Traditional models have harm categories hardcoded into their weights, requiring retraining when targeting new deployment contexts.
What hardware is required to run Shieldstral?
The model runs on a single 16GB NVIDIA GPU. Its 3-billion parameter size makes it computationally efficient compared to larger moderation models.
How was Shieldstral trained to distinguish different safety policies?
The model was trained on synthetic contrastive pairs created by an LLM specifically engineered to violate one policy while not violating a similar variant. This teaches the model to distinguish precise policy boundaries rather than simply memorizing categories.
- LFM2.5-2.6B: small and capable local AI model — liquid.ai 75 % match
- An Alien Mind: OpenAI's Chief Scientist on Intelligence We Don't Fully Understand — openai.com 74 % match
- Harness Engineering for Self-Improvement — lilianweng.github.io 72 % match
- Mistral3
- Shieldstral
- NVIDIA8
- Open Secure AI Alliance
- Forge
- Hugging Face12