新颖主张还是既视感?重新思考多模态自动事实核查中的“无污染”动态评估
Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking
July 26, 2026
作者: Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, Yupeng Li
cs.AI
摘要
多模态自动事实核查(MAFC)通过检索外部证据并进行推理来验证声明。然而,现有大多数静态基准存在污染风险:它们主要由可通过大语言模型内部知识(无需外部证据)验证的过时声明组成。这可能导致性能评估虚高,且无法真实反映模型在处理需要最新信息的新型声明时的能力。为此,新兴动态基准专门收集在大语言模型知识截止日期之后发布的声明,并假定其未被污染。本研究重新审视了这一假设,通过实验探究当前最先进的静态AVerTeC基准与我们新构建的动态ClaimReview2025Q4基准中存在的污染风险,及其对MAFC评估的影响。实验获得16项发现,其中三个关键结果为:(1) 动态评估虽能降低但无法消除污染风险——17.09%–29.30%的截止日期后声明仍可能受污染;(2) 大量新发布声明可直接验证,或通过综合截止日期前已有的多条公开知识进行验证;(3) 污染会显著导致MAFC性能虚高,使宏F1分数最高提升11.34个百分点,并扭曲系统排名。基于这些发现,我们在严格污染控制条件下重新评估了当前最优的大语言模型。本研究为可信的MAFC评估提供了实用指南。
English
Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to-date information. To address this, emerging dynamic benchmarks collect claims published after LLMs' knowledge cut-off dates, assuming they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the-art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlighting three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09\%--29.30\% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation.