SafeAtlas-VL:超越二元多模态安全——基于大规模数据与防护模型
SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models
August 29, 2026
作者: Zongrui Wang, Xiangyang Zhu, Sicheng Wang, Han Wang, Dingyi Rong, Zeyu Zhang, Chunyi Li, Yue Shi, Kaiwei Zhang, Zicheng Zhang, Yuan Tian, Qi Jia, Yan Teng, Wei Sun, Ning Liu, Guangtao Zhai
cs.AI
摘要
多模态安全审核需要区分源自视觉内容、用户意图和助手行为的风险。然而,现有的安全机制通常针对单一判断目标进行训练,并将安全评估简化为二元决策。因此,在多模态交互中,风险难以进行比较,模糊案例也被掩盖。我们提出了 SafeAtlas-VL,一个包含 150 万训练实例的数据集,将图像级、请求级和响应级判断置于五级有序尺度上。我们汇集了来自真实世界和合成来源的广泛安全相关数据,并应用了感知分歧的标注流程。生成的数据集涵盖 15 个危害类别和 55 个细分子类别,覆盖了广泛的多模态安全场景。我们还构建了 SafeAtlas-Bench,一个包含 5,000 个保留实例的测试集,用于评估五级预测和连续风险分数。在该数据集上,我们通过目标条件微调训练了 SafeAtlas Guard 系列模型,用于多模态安全检测。我们的模型不仅执行安全等级的五分类,还通过软累积有序头将安全映射为连续分数。实验结果表明,在我们的数据集上训练的防护模型展现出强大的泛化能力:即使不使用其他基准的训练集,它们也能在相应测试集上取得有竞争力的表现。值得注意的是,我们的 8B 模型取得了整体最佳性能,F1 分数比之前的 SOTA 高出约 4%。代码、数据和模型均已发布,以支持进一步研究。警告:本文包含的示例数据可能具有冒犯性、有害性、露骨或令人不安的内容。
English
Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for a single judgment target and reduce safety assessment to a binary decision. Consequently, risk becomes difficult to compare across a multimodal interaction, and ambiguous cases are obscured. We introduce SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale. We curate a broad collection of safety-relevant data from both real-world and synthetic sources and apply a disagreement-aware annotation procedure. The resulting dataset spans 15 harm categories and 55 fine-grained subcategories, covering a broad range of multimodal safety scenarios. We also construct SafeAtlas-Bench, a held-out set of 5,000 instances for evaluating five-level predictions and continuous risk scores. Upon this dataset, we train the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection. Our models not only perform five-way classification of safety levels but also map safety to continuous scores through a soft cumulative ordinal head. Experimental results demonstrate that guard models trained on our dataset exhibit strong generalization: even without using the training sets of other benchmarks, they achieve competitive performance on the corresponding test sets. Notably, our 8B model attains the overall best performance, outperforming the previous SOTA by approximately 4% in F1 score. Code, data, and models are released to support further research. Warning: this paper contains example data that may be offensive, harmful, graphic, or disturbing.