Mistral has released Shieldstral, an open-weight model designed to assess the safety of text and images. Unlike conventional safety systems, Shieldstral does not rely on predefined categories for harmful or prohibited content. Instead, a safety rule can be formulated as a simple yes-or-no question. These questions can be adjusted flexibly, while the context of the content can also be taken into account more effectively.
This means the same model can be adapted to different policies without having to retrain it. Mistral says Shieldstral delivers the same level of safety performance as much larger models. The model has 3 billion parameters and is available for download on Hugging Face under the Apache 2.0 license.