TECH FLOW Svět Androida
← Back to the stream
mistral.ai · picked by Petr Mišák · 60d ago

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

Source preview: Mistral's Shieldstral: 3B open-weights model for multimodal moderation
AI summary

Mistral released Shieldstral, a 3B open-weights model for text and image safety classification that operates as a question-answering system accepting plain-language policies without retraining. The model matches performance with models up to 7 times larger, runs on a single GPU, and is available under Apache 2.0 license.

The summary is written by AI from the source; it isn’t the newsroom’s opinion. For details, read the source.

11 people have already opened the source

AI questions & answers
How does Shieldstral differ from traditional content moderation models?

Shieldstral accepts moderation policies as natural language text at inference time, enabling adaptation to different contexts without retraining. Traditional models have harm categories hardcoded into their weights, requiring retraining when targeting new deployment contexts.

What hardware is required to run Shieldstral?

The model runs on a single 16GB NVIDIA GPU. Its 3-billion parameter size makes it computationally efficient compared to larger moderation models.

How was Shieldstral trained to distinguish different safety policies?

The model was trained on synthetic contrastive pairs created by an LLM specifically engineered to violate one policy while not violating a similar variant. This teaches the model to distinguish precise policy boundaries rather than simply memorizing categories.

Questions and answers are written by AI about the topic, not taken from the source; they aren’t the newsroom’s opinion.
Related from the stream
Mentions