星盾
Shieldstral
July 28, 2026
作者: Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli, Guillaume Lample, Maarten Buyl, Maximilian Augustin, Maximilian Müller, Pierre Stock, Tom Bewley, Wassim Bouaziz, Yimu Pan
cs.AI
摘要
我们推出了Shieldstral,这是一个30亿参数的政策自适应多模态安全分类器,在文本安全基准测试中能够匹配甚至超越近七倍于其体量的模型,并在多模态安全分类领域树立了新的行业标杆。Shieldstral将内容审核重构为二元问答任务,这种简洁的表述将多样化的审核任务统一为简单的是/否问题,使得原本具有不同分类体系的异构安全数据集能够在同一训练框架下实现整合。我们详细阐述了数据构建方案,涵盖约5410万个样本的整理与生成流程,以及用于评估政策适应性的细粒度评估数据集。这些要素共同使这个小型自适应模型得以匹敌甚至超越规模更大的模型。
English
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.