새로운 주장인가, 데자뷰인가? 멀티모달 자동 사실 검증을 위한 '오염 없는' 동적 평가 재고
Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking
July 26, 2026
저자: Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, Yupeng Li
cs.AI
초록
다중 모드 자동 팩트체킹(MAFC)은 외부 증거를 검색하고 추론하여 주장을 검증한다. 그러나 기존의 대부분의 정적 벤치마크는 오염 위험에 노출되어 있다. 즉, 주로 LLM의 내부 지식만으로 외부 증거 없이 검증 가능한 오래된 주장들로 구성되어 있어, 성능 추정치가 부풀려질 수 있으며 최신 정보를 필요로 하는 새로운 주장에 대한 실제 능력을 반영하지 못할 수 있다. 이를 해결하기 위해, 최근 등장한 동적 벤치마크는 LLM의 지식 마감일 이후에 게시된 주장들을 수집하여 오염되지 않았다고 가정한다. 본 연구는 최첨단(SOTA) 정적 벤치마크인 AVeriTeC와 새로 구축한 동적 벤치마크인 ClaimReview2025Q4에서 오염 위험을 실증적으로 조사하고, 이것이 MAFC 평가에 미치는 영향을 분석함으로써 이러한 가정을 재검토한다. 실험을 통해 16가지 결과를 도출하였으며, 다음 세 가지 주요 결과를 강조한다: (1) 동적 평가는 오염 위험을 줄이지만 완전히 제거하지는 않으며, 마감일 이후 주장의 17.09%~29.30%가 여전히 잠재적으로 오염될 수 있다. (2) 새로 게시된 많은 주장들은 마감일 이전에 공개된 여러 공개 지식을 직접 활용하거나 종합하여 검증할 수 있다. (3) 오염은 MAFC 성능에 통계적으로 유의미한 부풀림을 유발하여 Macro-F1을 최대 11.34포인트 증가시키고 시스템 순위를 왜곡할 수 있다. 이러한 발견을 바탕으로, 엄격한 오염 통제 환경에서 최첨단 LLM을 재평가한다. 본 연구는 신뢰할 수 있는 MAFC 평가를 위한 실용적 지침을 제공한다.
English
Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to-date information. To address this, emerging dynamic benchmarks collect claims published after LLMs' knowledge cut-off dates, assuming they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the-art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlighting three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09\%--29.30\% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation.