← Latest papers
💬 NLP

Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset

This paper presents a comprehensive pipeline that leverages large language models to construct structured, multilingual, and multimodal fact-checking datasets in French and German by aggregating ClaimReview feeds, extracting evidence, and generating justifications to address the limitations of existing misinformation resources.

Original authors: Z. Melce Hüsünbeyi, Virginie Mouilleron, Leonie Uhling, Daniel Foppe, Tatjana Scheffler, Djamé Seddah

Published 2026-03-16
📖 4 min read☕ Coffee break read

Original authors: Z. Melce Hüsünbeyi, Virginie Mouilleron, Leonie Uhling, Daniel Foppe, Tatjana Scheffler, Djamé Seddah

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a giant, chaotic marketplace. In this marketplace, people are shouting out all sorts of claims—some are true, some are lies, and some are just half-truths. Sometimes these claims are just words, but often they come with pictures, videos, or memes that make them look even more convincing.

The problem is that the "fact-checkers" (the librarians of this marketplace) are overwhelmed. They have to check thousands of claims every day, digging through official documents, expert interviews, and video footage to figure out what's real. Meanwhile, computers (AI) are trying to help, but they've been playing a game with a broken rulebook: they mostly look at text, ignore the pictures, and don't really understand how a human fact-checker actually does their job.

This paper introduces a new, super-powered toolkit to fix this. Here is how it works, broken down into simple steps:

1. The "Recipe Book" (The Pipeline)

The researchers built a giant, automated assembly line (a pipeline) to collect fact-checking stories in French and German.

  • The Ingredients: They didn't just grab random headlines. They went to the source: the actual articles written by professional fact-checkers.
  • The Sorting: They took these messy articles and organized them into a neat, structured format. They didn't just say "True" or "False." They broke down the proof into specific buckets, like a chef organizing ingredients:
    • Expert Testimony: What did the doctor or scientist say?
    • Numbers & Stats: What do the charts and graphs say?
    • Official Records: What do the government laws or court documents say?
    • Eyewitnesses: What did the person who was actually there say?
    • Multimedia: What do the photos and videos prove?

2. The "Translator" (The AI's New Job)

Once they had this organized "recipe book," they taught advanced AI models (Large Language Models) how to read it.

  • Task A (Evidence Extraction): Imagine giving the AI a messy pile of papers and asking it to pull out only the specific quotes and numbers that prove a point. The AI learned to sort these into the correct buckets (e.g., "This is an expert quote," "This is a statistic").
  • Task B (Justification Generation): After finding the proof, the AI had to write a short, clear explanation connecting the dots. It had to say, "Because the expert said X and the chart shows Y, this claim is False."

3. The "Taste Test" (Evaluation)

How do we know the AI is doing a good job? The researchers didn't just trust the computer.

  • They used a "Judge AI" (another super-smart AI) to grade the work on three things:
    1. Correctness: Did it make things up? (No hallucinations).
    2. Coherence: Did the explanation flow logically?
    3. Completeness: Did it miss any important proof?
  • They also had real humans grade the AI. The humans and the AI judges agreed: the best models were getting very good at mimicking how human fact-checkers think.

4. The "Multimodal" Magic

The biggest breakthrough here is that the AI isn't just reading text anymore. It's looking at images and videos too.

  • Think of a claim like "This video shows a politician at a rally."
  • An old AI might just read the text.
  • This new system looks at the video frames, reads the captions, checks the timestamp, and sees if the video actually matches the claim. It treats the video as a piece of evidence, just like a document.

Why Does This Matter?

Think of this paper as building a training school for AI fact-checkers.

  • Before: AI was like a student who only read the summary of a book and guessed the ending.
  • Now: AI is like a student who has been taught to read the whole book, check the footnotes, look at the photos, and write a detailed report on why the story is true or false.

The result is a massive, open library of fact-checked stories (in French and German) that teaches AI to be more transparent, more accurate, and better at spotting lies that hide behind pictures and videos. It's a step toward a future where we can trust our digital news a little bit more.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →