ChatPaper.aiChatPaper

汎化可能なディープフェイク動画検出のためのマルチエージェントフォレンジック推論

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

August 7, 2026
著者: Xuechao Zou, Shun Zhang, Kai Li, Yi Zhou, Xinyu Sun, Yuhui Chen, Zhe Wu, Congyan Lang, Junliang Xing
cs.AI

要旨

高度にリアルなディープフェイク動画を作成するための生成AIの悪意ある利用は、深刻な倫理的懸念を引き起こし、AI安全性に大きな課題を突きつけている。しかし、既存のディープフェイク動画ベンチマークは、最近の合成手法のカバー範囲が限られており、一般に信頼性の高い詳細なテキストアノテーションも欠いている。一方、従来の検出器やマルチモーダル大規模言語モデル(MLLM)は、単一モデルとして動作する場合も、単一の分析視点に依存する場合も、微細な偽造アーティファクトを捉え損ねることが多く、新興のAI生成手法への汎化が制限されている。これらの限界に対処するため、我々はFaceVid-Forensics-100Kを導入する。これは、10万本の動画からなり、Seedance 2.0などの最近の生成器を含む、フェイススワップ、フェイスリエナクトメント、顔全体の合成にわたる33種類の合成手法を網羅する大規模ディープフェイク動画データセットである。このデータセットは、高度なMLLMを活用したマルチモデル集約・競合解決パイプラインを通じて自動生成された、視覚的観察の詳細なテキストアノテーションと、判定結果と整合するフォレンジック説明を提供する。このベンチマークに基づき、我々は、テクスチャ、照明、動き、物理という4つの観点から偽造手がかりを独立に分析する4つの専門分野特化エージェントを用いるマルチエージェントフォレンジック推論フレームワークを提案する。次に、判定エージェントがそれらのレポートを統合し、説明とともに最終予測を生成する。ドメイン外テストセットに対する広範な評価により、我々のフレームワークは、完全に小規模なオープンソースMLLMのみで構成されているにもかかわらず、クローズドソースのGPTおよびGeminiモデルを含むすべての手法を上回り、このベンチマークの報告されたすべての指標で第1位を獲得した。プロジェクトページは https://xavierjiezou.github.io/ARGUS/ で入手可能である。
English
The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmarks provide limited coverage of recent synthesis methods and generally lack reliable fine-grained textual annotations. Meanwhile, conventional detectors and multimodal large language models (MLLMs), whether operating as a single model or relying on a single analytical perspective, often fail to capture subtle forgery artifacts, limiting their generalization to emerging AI-generated methods. To address these limitations, we introduce FaceVid-Forensics-100K, a large-scale deepfake video dataset comprising 100,000 videos and spanning 33 synthesis methods across face swapping, face reenactment, and entire-face synthesis, including recent generators such as Seedance 2.0. The dataset provides fine-grained textual annotations of visual observations and verdict-consistent forensic explanations, automatically synthesized through a multi-model aggregation and conflict-resolution pipeline powered by advanced MLLMs. Building on this benchmark, we propose a multi-agent forensic reasoning framework that employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives: texture, lighting, motion, and physics. A judge agent then reconciles their reports to produce a final prediction together with an explanation. Extensive evaluations on out-of-domain test sets show that, despite being composed entirely of small open-source MLLMs, our framework outperforms all methods including closed-source GPT and Gemini models and ranks first across all reported metrics on this benchmark. The project page is available at https://xavierjiezou.github.io/ARGUS/.