ChatPaper.aiChatPaper

우리는 실제 세계의 위기 사건에 대한 AI 생성 비디오 공격을 방어할 수 있는가? 탐지기, 생성기 및 사회적 확산에 대한 체계적 평가

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

August 14, 2026
저자: Shuo Liang, Yixing Ma, Pengfei Zhou, Xingyan Chen, Zihan Mei, Manting Li, Feihan Chen, Zhiwen Wang, Bin Xu, Haotian Zhang, Jiajun Song, Shiya Su, Run Liu, Zhenghang Ni, Yifa Yu, Jintao Hong, Bolong Feng, Yifei Liu, Zirui Zhang, Jingxuan Zhang, Songlin Zhao, Yifan Bai, Kang Tan, Yizhe Liu, Junhao Du, Yongtao Ge, Zhaopan Xv, Xinyuan Zhang, Mengru Ma, Chunhua Shen, Wei Wang, Yang You, Zheng Zhu, Kaipeng Zhang, Wangbo Zhao
cs.AI

초록

최근 비디오 생성기는 전쟁, 재난, 공공 비상사태 및 기타 실세계 위기에 대한 사실적인 묘사를 만들어낼 수 있으며, 이는 상당한 허위 정보 위험을 초래한다. 그러나 기존 벤치마크는 탐지 가능성이 생성 조건에 따라 어떻게 달라지는지, 사람들이 생성된 비디오를 어떻게 인식하는지, 탐지기가 사회적 확산 과정에서도 신뢰할 수 있는지 여부 등 이러한 상황에서의 탐지기와 생성기의 행동에 대한 제한적인 증거만 제공한다. 이러한 격차를 해소하기 위해, 우리는 실제 비디오를 앵커로 사용하는 AI 생성 비디오 탐지 벤치마크인 RA-Bench를 소개한다. RA-Bench는 17,886개의 비디오로 구성되며, 10개의 사회적 위험 범주에 걸친 1,830개의 실제 비디오 앵커와 4개의 오픈소스 및 5개의 클로즈드소스 생성기에서 생성된 16,056개의 생성 클립을 포함한다. RA-Bench를 기반으로, 우리는 평가를 세 가지 차원으로 구성한다. 먼저, 우리는 일곱 개의 전통적 탐지기, 세 가지 검토 설정에서의 열 개의 제로샷 멀티모달 모델, 그리고 AI 생성 비디오 탐지에 특별히 미세 조정된 두 개의 MLLM에 걸쳐 탐지기 일반화를 평가한다. 이러한 방법들 전반에 걸쳐, 세 탐지기 계열 중 어느 것도 RA-Bench 인스턴스에서 일관되게 일반화되지 않는다. 그런 다음 우리는 탐지 가능성이 생성 품질, 조건화 정보, 그리고 샘플링 시드에 따라 어떻게 달라지는지 조사한다. 이러한 분석은 생성 속성이 탐지기 계열에 서로 다르게 영향을 미치는 반면, 소스 수준의 탐지 패턴은 시드에 걸쳐 안정적으로 유지됨을 보여준다. 마지막으로, 우리는 사회적 확산 과정에서 인간의 진위성 판단과 탐지기 신뢰성을 연구한다. 우리는 사람들을 오도하는 비디오가 현재의 탐지기에게도 탐지하기 어렵고, 사회적 확산이 탐지를 더 어렵게 만든다는 것을 발견한다. 종합하면, 이러한 발견은 현재의 방법들이 사실적인 AI 생성 비디오를 탐지하는 데 어려움을 겪고 있음을 보여주며, 진화하는 비디오 생성기에 강건한 탐지기의 필요성을 강조한다.
English
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable during social dissemination. To address this gap, we introduce RA-Bench, a benchmark for AI-generated video detection that uses Real videos as Anchors. RA-Bench contains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based on RA-Bench, we organize our evaluation along three dimensions. We first assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs specifically fine-tuned on AI-generated video detection. Across these methods, none of the three detector families generalizes consistently across RA-Bench instances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability during social dissemination. We find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.