新規主張か、それともデジャヴか?マルチモーダル自動ファクトチェックにおける「汚染のない」動的評価の再考
Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking
July 26, 2026
著者: Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, Yupeng Li
cs.AI
要旨
多モーダル自動ファクトチェック(MAFC)は、外部証拠を検索し推論することで主張を検証する。しかし、既存の静的なベンチマークの大半は汚染のリスクを抱えている。すなわち、それらは主に、外部証拠を用いずにLLMの内部知識だけで検証可能な時代遅れの主張から構成されている。これにより性能評価が過大になる可能性があり、最新情報を必要とする新規の主張に対する真の能力を反映できない。この問題に対処するため、新たに登場した動的ベンチマークは、LLMの知識カットオフ日以降に公開された主張を収集し、汚染されていないと仮定している。本研究では、この仮定を再検討し、最先端(SOTA)の静的ベンチマークAVeriTeCと、新たに構築した動的ベンチマークClaimReview2025Q4の両方における汚染リスクと、それらがMAFC評価に与える影響を実証的に調査する。実験から16の知見が得られ、以下の3つの主要な結果が明らかになった。(1) 動的評価は汚染リスクを低減するが完全には排除せず、カットオフ日以降の主張の17.09%~29.30%が依然として汚染されている可能性がある。(2) 新たに公開された主張の多くは、カットオフ日以前に利用可能な複数の公開知識を合成するか、直接的に検証できる。(3) 汚染はMAFC性能に統計的に有意な過大評価を引き起こし、Macro-F1を最大11.34ポイント押し上げ、システムの順位を歪める可能性がある。これらの知見を踏まえ、厳密に汚染を制御した条件下でSOTA LLMを再評価する。本研究は、信頼性の高いMAFC評価のための実践的ガイドラインを提供する。
English
Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to-date information. To address this, emerging dynamic benchmarks collect claims published after LLMs' knowledge cut-off dates, assuming they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the-art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlighting three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09\%--29.30\% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation.