実世界の危機事象に対するAI生成ビデオ攻撃を防御できるか? 検出器・生成器・社会的拡散の系統的評価
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
August 14, 2026
著者: Shuo Liang, Yixing Ma, Pengfei Zhou, Xingyan Chen, Zihan Mei, Manting Li, Feihan Chen, Zhiwen Wang, Bin Xu, Haotian Zhang, Jiajun Song, Shiya Su, Run Liu, Zhenghang Ni, Yifa Yu, Jintao Hong, Bolong Feng, Yifei Liu, Zirui Zhang, Jingxuan Zhang, Songlin Zhao, Yifan Bai, Kang Tan, Yizhe Liu, Junhao Du, Yongtao Ge, Zhaopan Xv, Xinyuan Zhang, Mengru Ma, Chunhua Shen, Wei Wang, Yang You, Zheng Zhu, Kaipeng Zhang, Wangbo Zhao
cs.AI
要旨
最近のビデオ生成器は、戦争、災害、公衆緊急事態、その他の現実世界の危機に関する現実的な描写を作り出すことができ、誤情報の重大なリスクを生み出している。しかし、既存のベンチマークは、検出可能性が生成条件によってどのように変化するか、生成されたビデオを人々がどのように知覚するか、社会的拡散の際に検出器が信頼できるままであるかなど、このような状況における検出器と生成器の振る舞いに関する限られた知見しか提供していない。このギャップに対処するため、我々は実動画をアンカーとして使用するAI生成ビデオ検出のためのベンチマークであるRA-Benchを紹介する。RA-Benchは17,886本のビデオを含み、10の社会的リスクカテゴリにわたる1,830本の実動画アンカーと、4つのオープンソースおよび5つのクローズドソース生成器からの16,056本の生成クリップで構成される。RA-Benchに基づき、我々は評価を3つの側面に沿って構成する。まず、7つの従来型検出器、3つの評価設定における10のゼロショットマルチモーダルモデル、およびAI生成ビデオ検出に特化してファインチューニングされた2つのMLLMにわたって検出器の汎化を評価する。これらの手法全体で、3つの検出器ファミリーのいずれもRA-Benchのインスタンス全体で一貫して汎化しない。次に、検出可能性が生成品質、条件付け情報、サンプリングシードによってどのように変化するかを調べる。これらの分析は、生成特性が検出器ファミリーに異なる影響を与える一方、生成元レベルの検出パターンはシード間で安定していることを示している。最後に、人間の真偽判断と社会的拡散中の検出器の信頼性を研究する。人々を誤解させるビデオは現在の検出器にとっても困難であり、社会的拡散によって検出がより困難になることを我々は見いだす。これらの発見は、現在の手法が現実的なAI生成ビデオの検出に苦慮していることを示しており、進化するビデオ生成器に対して頑健な検出器の必要性を強調している。
English
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable during social dissemination. To address this gap, we introduce RA-Bench, a benchmark for AI-generated video detection that uses Real videos as Anchors. RA-Bench contains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based on RA-Bench, we organize our evaluation along three dimensions. We first assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs specifically fine-tuned on AI-generated video detection. Across these methods, none of the three detector families generalizes consistently across RA-Bench instances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability during social dissemination. We find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.