ChatPaper.aiChatPaper

실드스트럴

Shieldstral

July 28, 2026
저자: Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli, Guillaume Lample, Maarten Buyl, Maximilian Augustin, Maximilian Müller, Pierre Stock, Tom Bewley, Wassim Bouaziz, Yimu Pan
cs.AI

초록

우리는 텍스트 안전 벤치마크에서 약 7배 더 큰 모델들과 동등하거나 능가하는 성능을 보이며, 멀티모달 안전 분류 분야에서 새로운 최고 수준을 달성한 3B 파라미터 규모의 정책 적응형 멀티모달 안전 분류기 Shieldstral을 소개합니다. Shieldstral은 콘텐츠 조정을 이진 질문-응답 과제로 정식화합니다. 이 간단한 정식화는 다양한 조정 작업을 단일 예/아니오 문제로 통합하여, 상이한 분류 체계를 가진 이질적인 안전 데이터셋을 하나의 훈련 프레임워크 아래 통합할 수 있게 합니다. 우리는 약 5410만 개의 샘플에 대한 선별 및 생성 과정을 포함한 데이터 구축 방안과, 정책 적응성을 평가하기 위한 세분화된 평가 세트를 제시합니다. 이러한 요소들을 통해 작은 규모의 적응형 모델이 훨씬 더 큰 모델들과 동등하거나 능가하는 성능을 발휘할 수 있습니다.
English
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.