← Latest papers
💬 NLP

RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild

This paper introduces RW-Post, a new auditable benchmark for real-world multimodal fact-checking that links social media posts with human-verified evidence and reasoning traces, alongside the AgentFact baseline, to reveal significant gaps in current models' ability to faithfully ground visual claims in evidence.

Original authors: Danni Xu, Shaojing Fan, Harry Cheng, Mohan Kankanhalli

Published 2026-05-13
📖 5 min read🧠 Deep dive

Original authors: Danni Xu, Shaojing Fan, Harry Cheng, Mohan Kankanhalli

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Fake News" Trap

Imagine you are scrolling through social media and see a post with a shocking headline and a dramatic photo. The text says, "This is a nuclear missile being moved to a border!" and the photo shows a truck with a long object.

The problem is that the photo might be real, but the story is a lie. Maybe the truck is actually carrying a parade float from 2015, not a missile. Current AI tools are getting better at spotting these lies, but they often make mistakes in two ways:

  1. They guess: They say "This looks fake" without showing their work.
  2. They hallucinate: They make up fake reasons or cite evidence that doesn't actually exist, just to sound convincing.

The Solution: RW-Post (The "Detective's Notebook")

The authors introduce a new tool called RW-Post. Think of this not just as a test, but as a gold-standard training manual for AI detectives.

In the real world, when a human fact-checker investigates a rumor, they don't just guess. They:

  • Find the original post (to see the context).
  • Gather proof (like old photos, news articles, or videos).
  • Write down their step-by-step logic (e.g., "The truck in the photo matches a 2015 parade, not a 2023 missile").

RW-Post is a massive dataset that does exactly this for AI. It pairs real-world social media posts with:

  • The original claim.
  • Reasoning traces: A step-by-step log of how a human figured out the truth.
  • Explicitly linked evidence: Specific links to the exact photos or articles that prove the claim true or false.

It's like giving an AI student a textbook where every answer is backed up by a specific page number and a direct quote, rather than just giving them a multiple-choice quiz.

The Three Ways to Take the Test

The paper tests AI models in three different "exam modes" to see where they struggle:

  1. Closed-Book (The Memory Test): The AI sees the post but has to rely only on what it already knows from its training. It can't look anything up.
    • Result: AI struggles here. It often guesses or makes things up.
  2. Evidence-Bounded (The Open-Book Test): The AI sees the post and is handed a specific folder of the correct evidence (the exact articles and photos the human fact-checker used).
    • Result: AI gets much smarter! When it has the right facts, it can solve the puzzle.
  3. Open-Web (The Real-World Hunt): The AI has to go out onto the internet, search for the evidence itself, and then solve the puzzle.
    • Result: This is the hardest mode. The AI often gets distracted or picks the wrong links.

The New AI Detective: AgentFact

To test these models, the authors built a reference system called AgentFact. Imagine this as a team of specialized detectives working together, rather than one person trying to do everything.

  • The Planner: Decides what questions to ask.
  • The Text Hunter: Searches the web for articles.
  • The Image Hunter: Does a "reverse image search" to see if the photo has been used before or edited.
  • The Reasoner: Looks at all the clues and decides if the story is true or false.
  • The Reporter: Writes the final explanation, making sure to cite exactly which clue proved what.

What They Found (The Results)

The experiments revealed some interesting truths about current AI:

  • Evidence is King: When AI models were given the evidence (the "Open-Book" test), their accuracy jumped significantly. This proves that AI isn't "smart" enough to solve these mysteries on its own; it needs the facts.
  • The "Visual" Blind Spot: Even when AI models were given both the text and the evidence, they still struggled to understand the images. They often missed visual clues, like realizing a photo was from a different year or location.
  • The "Fake Proof" Problem: When AI tries to explain why it thinks something is fake, it often makes up reasons that sound good but aren't supported by the actual evidence. The new dataset (RW-Post) helps catch this because it forces the AI to link its reasoning to real, auditable proof.

The Bottom Line

The paper argues that to stop misinformation, we can't just rely on AI guessing. We need systems that act like rigorous journalists: they must find the original source, gather hard evidence, and show their work step-by-step.

RW-Post provides the training ground to teach AI how to do this, and AgentFact shows us what a well-organized, evidence-based AI detective looks like. The main takeaway is that while AI is getting better, it still needs a lot of help finding and understanding the visual and textual evidence to tell the truth.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →