ChatPaper.aiChatPaper

星盾

Shieldstral

July 28, 2026
作者: Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli, Guillaume Lample, Maarten Buyl, Maximilian Augustin, Maximilian Müller, Pierre Stock, Tom Bewley, Wassim Bouaziz, Yimu Pan
cs.AI

摘要

我們推出 Shieldstral,這是一個擁有 30 億參數的政策自適應多模態安全分類器。在文本安全基準測試中,其表現與規模近七倍的模型相當或更優,並在多模態安全分類領域樹立了新的技術標杆。Shieldstral 將內容審核形式化為二元問答任務。這種簡潔的表述將多樣化的審核任務統整為單一的「是/否」問題,使具備不同分類架構的異質安全資料集能在同一訓練框架下整合。我們詳述資料建構方法,涵蓋約 5,410 萬筆樣本的策劃與生成,以及用於評估政策適應性的細粒度評估集。這些要素共同使一個小型自適應模型能與規模大得多的模型匹敵甚至超越。
English
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.