SafeAtlas-VL:以大規模資料與防護模型超越二元多模態安全
SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models
August 29, 2026
作者: Zongrui Wang, Xiangyang Zhu, Sicheng Wang, Han Wang, Dingyi Rong, Zeyu Zhang, Chunyi Li, Yue Shi, Kaiwei Zhang, Zicheng Zhang, Yuan Tian, Qi Jia, Yan Teng, Wei Sun, Ning Liu, Guangtao Zhai
cs.AI
摘要
多模態安全審核需要區分來自視覺內容、使用者意圖與助理行為所產生的風險。然而,現有的安全防護機制通常針對單一判斷目標進行訓練,並將安全評估簡化為二元決策。因此,風險在多模態互動中難以比較,模糊案例亦被掩蓋。我們提出 SafeAtlas-VL,一個包含 150 萬筆訓練實例的資料集,將影像層級、請求層級與回應層級的判斷置於五級有序量表上。我們從真實世界與合成來源中蒐集廣泛的安全相關資料,並採用分歧感知的標註程序。由此產生的資料集涵蓋 15 個危害類別與 55 個細粒度子類別,覆蓋廣泛的多模態安全情境。我們亦建構 SafeAtlas-Bench,一個包含 5,000 個實例的保留測試集,用於評估五級預測與連續風險分數。在此資料集上,我們透過目標條件微調訓練 SafeAtlas Guard 系列模型,以進行多模態安全偵測。我們的模型不僅執行安全等級的五向分類,也透過軟性累積序數輸出頭將安全映射為連續分數。實驗結果顯示,在我們的資料集上訓練的防護模型展現出強大的泛化能力:即使不使用其他基準的訓練集,它們在對應的測試集上也能達到具競爭力的表現。值得注意的是,我們的 8B 模型取得了整體最佳表現,在 F1 分數上比先前的 SOTA 高出約 4%。我們已釋出程式碼、資料與模型,以支持後續研究。警告:本文包含可能具有冒犯性、有害、露骨或令人不安的示例資料。
English
Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for a single judgment target and reduce safety assessment to a binary decision. Consequently, risk becomes difficult to compare across a multimodal interaction, and ambiguous cases are obscured. We introduce SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale. We curate a broad collection of safety-relevant data from both real-world and synthetic sources and apply a disagreement-aware annotation procedure. The resulting dataset spans 15 harm categories and 55 fine-grained subcategories, covering a broad range of multimodal safety scenarios. We also construct SafeAtlas-Bench, a held-out set of 5,000 instances for evaluating five-level predictions and continuous risk scores. Upon this dataset, we train the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection. Our models not only perform five-way classification of safety levels but also map safety to continuous scores through a soft cumulative ordinal head. Experimental results demonstrate that guard models trained on our dataset exhibit strong generalization: even without using the training sets of other benchmarks, they achieve competitive performance on the corresponding test sets. Notably, our 8B model attains the overall best performance, outperforming the previous SOTA by approximately 4% in F1 score. Code, data, and models are released to support further research. Warning: this paper contains example data that may be offensive, harmful, graphic, or disturbing.