シールドストラル
Shieldstral
July 28, 2026
著者: Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli, Guillaume Lample, Maarten Buyl, Maximilian Augustin, Maximilian Müller, Pierre Stock, Tom Bewley, Wassim Bouaziz, Yimu Pan
cs.AI
要旨
本稿では、3Bパラメータのポリシー適応型マルチモーダル安全性分類器であるShieldstralを紹介する。本モデルは、テキスト安全性ベンチマークにおいて約7倍の規模のモデルと同等以上の性能を示し、マルチモーダル安全性分類において新たな最先端を達成する。Shieldstralは、コンテンツモデレーションを二値質問応答タスクとして定式化する。この単純な定式化により、多様なモデレーションタスクが単一のyes/no問題に統合され、異なる分類体系を持つ異種の安全性データセットを一つの学習フレームワークのもとで統合することを可能にする。本稿では、約5410万サンプルのキュレーションおよび生成、ならびにポリシー適応性を評価するための詳細な評価セットを含むデータ構築手法を提示する。これらの要素により、小規模な適応型モデルがはるかに大規模なモデルと同等以上の性能を達成することが可能となる。
English
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.