ChatPaper.aiChatPaper

SafeAtlas-VL: 大規模データとガードモデルによる二値分類を超えたマルチモーダル安全性

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

August 29, 2026
著者: Zongrui Wang, Xiangyang Zhu, Sicheng Wang, Han Wang, Dingyi Rong, Zeyu Zhang, Chunyi Li, Yue Shi, Kaiwei Zhang, Zicheng Zhang, Yuan Tian, Qi Jia, Yan Teng, Wei Sun, Ning Liu, Guangtao Zhai
cs.AI

要旨

マルチモーダル安全モデレーションでは、視覚コンテンツ、ユーザーの意図、およびアシスタントの動作から生じるリスクを区別する必要があります。しかし、既存のセーフガードは通常、単一の判定対象に対して訓練されており、安全性評価を二値判定に還元しています。その結果、マルチモーダルな対話全体でリスクを比較することが難しく、曖昧なケースが不明瞭なままになっています。我々は、画像レベル、リクエストレベル、応答レベルの判定を5段階の順序スケールで評価する、150万件の訓練インスタンスからなるデータセットSafeAtlas-VLを導入します。実世界と合成の両方のソースから安全性に関連する広範なデータを収集し、不一致を考慮したアノテーション手順を適用します。結果として得られたデータセットは、15の危害カテゴリと55の細粒度サブカテゴリにわたり、広範囲のマルチモーダル安全シナリオをカバーしています。また、5段階の予測と連続的なリスクスコアを評価するための5,000インスタンスのホールドアウトセットであるSafeAtlas-Benchも構築します。このデータセットに基づき、マルチモーダル安全検出のためにターゲット条件付きチューニングを用いてSafeAtlas Guardシリーズのモデルを訓練します。我々のモデルは、安全性レベルの5クラス分類を実行するだけでなく、ソフト累積順序ヘッドを通じて安全性を連続スコアにマッピングします。実験結果は、我々のデータセットで訓練されたガードモデルが強い汎化を示すことを実証しています。他のベンチマークの訓練セットを使用しなくても、対応するテストセットで競争力のある性能を達成します。注目すべきことに、我々の8Bモデルは全体で最高の性能を達成し、F1スコアにおいて従来のSOTAを約4%上回りました。さらなる研究を支援するため、コード、データ、モデルを公開しています。注意: 本論文には、不快、有害、露骨、または動揺を与える可能性のあるサンプルデータが含まれています。
English
Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for a single judgment target and reduce safety assessment to a binary decision. Consequently, risk becomes difficult to compare across a multimodal interaction, and ambiguous cases are obscured. We introduce SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale. We curate a broad collection of safety-relevant data from both real-world and synthetic sources and apply a disagreement-aware annotation procedure. The resulting dataset spans 15 harm categories and 55 fine-grained subcategories, covering a broad range of multimodal safety scenarios. We also construct SafeAtlas-Bench, a held-out set of 5,000 instances for evaluating five-level predictions and continuous risk scores. Upon this dataset, we train the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection. Our models not only perform five-way classification of safety levels but also map safety to continuous scores through a soft cumulative ordinal head. Experimental results demonstrate that guard models trained on our dataset exhibit strong generalization: even without using the training sets of other benchmarks, they achieve competitive performance on the corresponding test sets. Notably, our 8B model attains the overall best performance, outperforming the previous SOTA by approximately 4% in F1 score. Code, data, and models are released to support further research. Warning: this paper contains example data that may be offensive, harmful, graphic, or disturbing.