ChatPaper.aiChatPaper

我们能否防御针对现实世界危机事件的AI生成视频攻击?——检测器、生成器与社交传播的系统性评估

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

August 14, 2026
作者: Shuo Liang, Yixing Ma, Pengfei Zhou, Xingyan Chen, Zihan Mei, Manting Li, Feihan Chen, Zhiwen Wang, Bin Xu, Haotian Zhang, Jiajun Song, Shiya Su, Run Liu, Zhenghang Ni, Yifa Yu, Jintao Hong, Bolong Feng, Yifei Liu, Zirui Zhang, Jingxuan Zhang, Songlin Zhao, Yifan Bai, Kang Tan, Yizhe Liu, Junhao Du, Yongtao Ge, Zhaopan Xv, Xinyuan Zhang, Mengru Ma, Chunhua Shen, Wei Wang, Yang You, Zheng Zhu, Kaipeng Zhang, Wangbo Zhao
cs.AI

摘要

近期,视频生成器能够编造出战争、灾难、公共突发事件及其他真实世界危机的逼真场景,从而造成重大的虚假信息风险。然而,现有基准测试在此类情境下提供的关于检测器和生成器行为的证据有限,包括可检测性如何随生成条件变化、人们如何感知生成的视频,以及检测器在社交传播过程中是否依然可靠。为填补这一空白,我们提出了RA-Bench——一个以真实视频为锚点(Real videos as Anchors)的AI生成视频检测基准。RA-Bench包含17,886个视频,涵盖10个社会风险类别中的1,830个真实视频锚点,以及来自四个开源和五个闭源生成器的16,056个生成片段。基于RA-Bench,我们从三个维度组织评估。首先,我们评估了七个传统检测器、三种评测设置下的十个零样本多模态模型,以及两个专门针对AI生成视频检测微调的多模态大语言模型(MLLM)的检测器泛化能力。在这些方法中,三个检测器家族均未能在RA-Bench实例上实现一致的泛化。随后,我们考察了可检测性如何随生成质量、条件信息和采样种子变化。这些分析表明,生成属性对不同检测器家族的影响各不相同,而源级检测模式在不同种子间保持稳定。最后,我们研究了社交传播中人类的真实性判断与检测器可靠性。我们发现,能够误导人类的视频对现有检测器而言同样难以检测,且社交传播会进一步增加检测难度。综合来看,这些发现表明当前方法难以检测逼真的AI生成视频,凸显了对能够应对不断演进的视频生成器的鲁棒检测器的需求。
English
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable during social dissemination. To address this gap, we introduce RA-Bench, a benchmark for AI-generated video detection that uses Real videos as Anchors. RA-Bench contains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based on RA-Bench, we organize our evaluation along three dimensions. We first assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs specifically fine-tuned on AI-generated video detection. Across these methods, none of the three detector families generalizes consistently across RA-Bench instances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability during social dissemination. We find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.