新穎主張或既視感?重新思考多模態自動事實查核中的「無污染」動態評估
Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking
July 26, 2026
作者: Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, Yupeng Li
cs.AI
摘要
多模態自動事實查核(MAFC)透過檢索並推理外部證據來驗證主張。然而,現有靜態基準大多存在污染風險:它們主要由過時的主張構成,這些主張僅需利用大型語言模型(LLM)的內在知識即可驗證,無需外部證據。這可能高估效能估算,無法真實反映模型在處理需要最新資訊的新穎主張時的能力。為解決此問題,新興的動態基準收集在LLM知識截止日期後發布的主張,假設其未被污染。本研究重新審視此假設,透過實驗探討最先進(SOTA)靜態基準AVeriTeC與我們新建構的動態基準ClaimReview2025Q4中的污染風險,及其對MAFC評估的影響。實驗得出16項發現,凸顯三大關鍵結果:(1)動態評估雖降低但未消除污染風險,截至截止日後仍有17.09%–29.30%的主張可能受污染;(2)許多新發布的主張可直接驗證,或透過綜合截止日前已公開的多項知識進行驗證;(3)污染可能導致MAFC效能出現統計顯著的高估,使宏觀F1分數最多提升11.34個百分點,並扭曲系統排名。基於這些發現,我們在嚴格控制的無污染環境下重新評估SOTA LLM。本研究為可信的MAFC評估提供實務指引。
English
Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to-date information. To address this, emerging dynamic benchmarks collect claims published after LLMs' knowledge cut-off dates, assuming they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the-art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlighting three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09\%--29.30\% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation.