← Latest papers
💬 NLP

Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking

This paper challenges the assumption that dynamic benchmarks are inherently contamination-free by demonstrating that a significant portion of post-cut-off claims remain verifiable via pre-existing knowledge, which can artificially inflate multimodal fact-checking performance and distort system rankings, thereby necessitating stricter evaluation protocols.

Original authors: Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, Yupeng Li

Published 2026-07-28✓ Author reviewed
📖 4 min read☕ Coffee break read

Original authors: Haorui He, Xinwen Chen, Dacheng Wen, Reynold Cheng, Francis C. M. Lau, Yupeng Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are training a super-smart robot to be a detective. Its job is to solve mysteries by finding clues on the internet, reading them, and deciding if a story is true or false. This is called "automated fact-checking." But here's the tricky part: to make sure the robot is actually learning to be a detective and not just a memorizer, we have to test it with new mysteries it has never seen before. If we give it a mystery it already knows the answer to because it read it in its training books, we can't tell if it's good at finding clues or just good at remembering. This paper dives into a specific corner of computer science where we try to build these detective robots for stories that mix text and pictures (like a viral meme or a news video). The big question is: Are we testing these robots fairly, or are we accidentally giving them the answers before the test even starts?

The researchers behind this paper, a team from Hong Kong and Beijing, decided to investigate a popular idea in the tech world. Many scientists thought that if they created "dynamic" tests—using news stories published after the robot's training stopped—they would be safe from cheating. They assumed that if a story was new, the robot couldn't possibly know the answer. But the authors suspected this might be a trap. They asked: What if the robot doesn't need to search for new clues because the answer is already hidden in its memory, even for "new" stories?

To find out, they built a special test lab. They took a brand-new set of fact-checking stories from late 2025 (after the robots' training cut-off) and compared them against an older set of stories. They used a clever method to see if the robots could write a fact-checking article using only their internal memory, without looking up anything new. If the robot's memory-only article looked a lot like the real, human-written fact-check, they flagged it as "contaminated"—meaning the robot was cheating by using old knowledge instead of doing the hard work of searching.

Their investigation revealed some surprising and slightly disappointing news. First, they found that even with these "fresh" dynamic tests, about 17% to 29% of the stories were still contaminated. It's like giving a student a test on a new topic, but the student realizes the answer is just a combination of three old facts they already memorized. For example, a story might be about a new rumor regarding a famous person's age; the robot doesn't need to search the web because it already knows the person's birth year and the legal age for a job, so it can solve the puzzle instantly.

The team also discovered that this "cheating" makes the robots look much smarter than they really are. When they tested the robots on the contaminated stories, the robots scored high. But when they switched to the truly fresh stories that the robots couldn't solve from memory, the scores dropped significantly—sometimes by as much as 11 points. It's the difference between a runner who looks fast because they are running on a treadmill they've practiced on, versus running on a new, bumpy trail. The contamination was so strong that it even changed the ranking of which robot was the "best." The robot that looked like the champion on the old test wasn't even the top performer on the fair, clean test.

In the end, the authors re-ran the tests with a strict rule: only use stories that no robot could solve from memory. Even with this fair playing field, the best robots still only got about 56% of the answers right. This suggests that while our detective robots are getting better, they still have a long way to go before they can truly handle the messy, fast-moving world of real-time news without getting confused or relying on old memories. The paper concludes that simply using "new" dates isn't enough to stop cheating; we need smarter ways to ensure the robots are actually doing the detective work, not just recalling the past.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →