ChatPaper.aiChatPaper

我們能否防禦針對真實世界危機事件的AI生成影片攻擊?對偵測器、生成器與社群傳播的系統性評估

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

August 14, 2026
作者: Shuo Liang, Yixing Ma, Pengfei Zhou, Xingyan Chen, Zihan Mei, Manting Li, Feihan Chen, Zhiwen Wang, Bin Xu, Haotian Zhang, Jiajun Song, Shiya Su, Run Liu, Zhenghang Ni, Yifa Yu, Jintao Hong, Bolong Feng, Yifei Liu, Zirui Zhang, Jingxuan Zhang, Songlin Zhao, Yifan Bai, Kang Tan, Yizhe Liu, Junhao Du, Yongtao Ge, Zhaopan Xv, Xinyuan Zhang, Mengru Ma, Chunhua Shen, Wei Wang, Yang You, Zheng Zhu, Kaipeng Zhang, Wangbo Zhao
cs.AI

摘要

近期的影片生成器能夠製造關於戰爭、災難、公共緊急事件及其他現實世界危機的逼真描繪,造成重大的錯誤資訊風險。然而,現有基準在此類情境下提供的證據有限,包括可檢測性如何隨生成條件變化、人們如何看待生成的影片,以及在社群傳播過程中檢測器是否仍然可靠。為填補這一缺口,我們提出RA-Bench,一個以真實影片為錨點(Anchors)的AI生成影片檢測基準。RA-Bench包含17,886部影片,涵蓋10個社會風險類別的1,830個真實影片錨點,以及來自四個開源與五個閉源生成器的16,056個生成片段。基於RA-Bench,我們沿三個維度組織評估。首先,我們評估七種傳統檢測器、三種審查情境下的十種零樣本多模態模型,以及兩種專門針對AI生成影片檢測進行微調的多模態大語言模型的檢測器泛化能力。在這些方法中,三個檢測器家族均未能在RA-Bench實例上持續泛化。接著,我們檢視可檢測性如何隨生成品質、條件資訊和取樣種子而變化。這些分析表明,生成屬性對不同檢測器家族的影響不同,而來源層級的檢測模式在不同種子間保持穩定。最後,我們研究社群傳播期間的人類真實性判斷與檢測器可靠性。我們發現,誤導人們的影片對當前檢測器而言同樣難以偵測,且社群傳播使檢測更加困難。綜合而言,這些發現表明當前方法難以檢測逼真的AI生成影片,凸顯了對能夠應對不斷演進的影片生成器之穩健檢測器的需求。
English
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable during social dissemination. To address this gap, we introduce RA-Bench, a benchmark for AI-generated video detection that uses Real videos as Anchors. RA-Bench contains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based on RA-Bench, we organize our evaluation along three dimensions. We first assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs specifically fine-tuned on AI-generated video detection. Across these methods, none of the three detector families generalizes consistently across RA-Bench instances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability during social dissemination. We find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.