Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking
Abstract
Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static benchmarks risk contamination: they primarily consist of outdated claims verifiable using an LLM's internal knowledge without external evidence. This can inflate performance estimates and fail to reflect true capability on novel claims that require up-to-date information. To address this, emerging dynamic benchmarks collect claims published after LLMs' knowledge cut-off dates, assuming they are uncontaminated. This work revisits this assumption by empirically studying contamination risks in both the state-of-the-art (SOTA) static AVeriTeC benchmark and our newly constructed dynamic ClaimReview2025Q4 benchmark, as well as their impact on MAFC evaluation. Our experiments yield 16 findings, highlighting three key results: (1) Dynamic evaluation reduces but does not eliminate contamination risks, as 17.09\%--29.30\% of post-cut-off claims remain potentially contaminated; (2) Many newly published claims can be verified either directly or by synthesizing multiple pieces of public knowledge available before the cut-off; and (3) Contamination can induce statistically significant inflation in MAFC performance, increasing Macro-F1 by up to 11.34 points and distorting system rankings. In light of these findings, we re-evaluate SOTA LLMs under a strictly contamination-controlled setting. Our study provides practical guidelines for trustworthy MAFC evaluation.
Community
“Novel Claim or Déjà Vu? Rethinking ‘Contamination-Free’ Dynamic Evaluation for Multimodal Automated Fact-Checking” (ACM MM 2026 Accepted)
🤔 High benchmark scores ≠ real-world fact-checking ability.
Static MAFC datasets let models take shortcuts via pretraining memory or data contamination.
🤯 Even “dynamic” benchmarks (post-cutoff claims) aren’t safe: many new claims can still be verified with pre-cutoff knowledge.
We introduce a Contamination Detection Pipeline that quantifies LLM/VLM knowledge contamination risk and removes tainted samples to reveal true performance on unseen claims.
💡 Key findings:
• New claims ≠ truly new
• Dynamic eval cuts risk but doesn’t eliminate it (17–29% still contaminated)
• Contamination can inflate Macro-F1 by up to 11.34 points and distort rankings
Let’s rethink how we evaluate real trustworthiness, not just leaderboard scores.
⭐️ Stars & discussions welcome!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Evidence Triangulation for Multimodal Fact-Checking in the Wild (2026)
- CREDENCE: Claim Reduction for Decomposition&Enhanced Credibility -- Semantic Metrics and Convergence Analysis (2026)
- Know Your Source: A Public Knowledge Store for Media Background Checks (2026)
- CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps (2026)
- Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting (2026)
- ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence (2026)
- HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.23514 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper