ChatPaper.aiChatPaper

SafeAtlas-VL: 대규모 데이터와 가드 모델을 통한 이진 분류를 넘어선 멀티모달 안전성

SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models

August 29, 2026
저자: Zongrui Wang, Xiangyang Zhu, Sicheng Wang, Han Wang, Dingyi Rong, Zeyu Zhang, Chunyi Li, Yue Shi, Kaiwei Zhang, Zicheng Zhang, Yuan Tian, Qi Jia, Yan Teng, Wei Sun, Ning Liu, Guangtao Zhai
cs.AI

초록

멀티모달 안전 모더레이션은 시각적 콘텐츠, 사용자 의도, 그리고 어시스턴트 행동에서 발생하는 위험을 구분해야 한다. 그러나 기존의 안전장치는 일반적으로 단일 판단 대상을 위해 훈련되며, 안전 평가를 이진 판단으로 축소한다. 결과적으로, 멀티모달 상호작용 전반에 걸쳐 위험을 비교하기 어려워지고, 모호한 사례는 가려지게 된다. 우리는 이미지 수준, 요청 수준, 응답 수준의 판단을 5단계 순서 척도로 배치한 150만 개의 훈련 인스턴스로 구성된 데이터셋 SafeAtlas-VL을 소개한다. 우리는 실제 세계와 합성 소스 모두에서 안전 관련 데이터를 광범위하게 수집하고, 불일치를 고려한 주석 절차를 적용한다. 결과 데이터셋은 15개의 유해 범주와 55개의 세분화된 하위 범주에 걸쳐 있으며, 광범위한 멀티모달 안전 시나리오를 포괄한다. 또한 5단계 예측과 연속 위험 점수를 평가하기 위한 5,000개 인스턴스의 분리된 평가 세트인 SafeAtlas-Bench를 구축한다. 이 데이터셋을 기반으로, 우리는 멀티모달 안전 탐지를 위한 타겟 조건부 튜닝을 통해 SafeAtlas Guard 시리즈 모델을 훈련한다. 우리의 모델은 안전 수준의 5-방향 분류를 수행할 뿐만 아니라, 소프트 누적 순서 헤드를 통해 안전을 연속 점수로 매핑한다. 실험 결과는 우리 데이터셋으로 훈련된 가드 모델이 강력한 일반화를 보임을 입증한다. 다른 벤치마크의 훈련 세트를 사용하지 않더라도 해당 테스트 세트에서 경쟁력 있는 성능을 달성한다. 특히, 우리의 8B 모델은 전반적으로 최고 성능을 달성하여, 이전 SOTA보다 F1 점수에서 약 4% 향상된 성능을 보인다. 코드, 데이터, 모델은 추가 연구를 지원하기 위해 공개된다. 경고: 본 논문은 공격적이거나, 유해하거나, 적나라하거나, 불쾌감을 줄 수 있는 예시 데이터를 포함하고 있다.
English
Multimodal safety moderation requires distinguishing risks arising from visual content, user intent, and assistant behavior. Existing safeguards, however, are typically trained for a single judgment target and reduce safety assessment to a binary decision. Consequently, risk becomes difficult to compare across a multimodal interaction, and ambiguous cases are obscured. We introduce SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale. We curate a broad collection of safety-relevant data from both real-world and synthetic sources and apply a disagreement-aware annotation procedure. The resulting dataset spans 15 harm categories and 55 fine-grained subcategories, covering a broad range of multimodal safety scenarios. We also construct SafeAtlas-Bench, a held-out set of 5,000 instances for evaluating five-level predictions and continuous risk scores. Upon this dataset, we train the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection. Our models not only perform five-way classification of safety levels but also map safety to continuous scores through a soft cumulative ordinal head. Experimental results demonstrate that guard models trained on our dataset exhibit strong generalization: even without using the training sets of other benchmarks, they achieve competitive performance on the corresponding test sets. Notably, our 8B model attains the overall best performance, outperforming the previous SOTA by approximately 4% in F1 score. Code, data, and models are released to support further research. Warning: this paper contains example data that may be offensive, harmful, graphic, or disturbing.