← Latest papers
💬 NLP

MEVER: Multi-Modal and Explainable Claim Verification with Graph-based Evidence Retrieval

MEVER is a novel multi-modal framework that performs joint evidence retrieval, claim verification, and explanation generation using a two-layer multi-modal graph and a Fusion-in-Decoder architecture, supported by a newly introduced scientific dataset called AIChartClaim.

Original authors: Delvin Ce Zhang, Suhan Cui, Zhelin Chu, Xianren Zhang, Dongwon Lee

Published 2026-02-11
📖 3 min read☕ Coffee break read

Original authors: Delvin Ce Zhang, Suhan Cui, Zhelin Chu, Xianren Zhang, Dongwon Lee

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a judge in a high-stakes courtroom, but instead of just listening to witnesses speak, you also have to look at complex charts, graphs, and scientific diagrams to decide if someone is telling the truth.

The problem is, most "AI judges" today are a bit nearsighted. They can read the written testimony (text), but they are often "blind" to the evidence shown in a picture (visuals). Even worse, when they do make a decision, they often can't explain why they reached that conclusion, leaving everyone in the courtroom confused.

This paper introduces MEVER, a new kind of AI judge designed to be both multi-modal (it can see and read) and explainable (it can tell you its reasoning).

Here is how MEVER works, broken down into three simple steps:

1. The Super-Powered Librarian (Evidence Retrieval)

Before a trial starts, a judge needs to find the right files. Most AI models just search through text files. MEVER, however, uses a "Two-Layer Multi-Modal Graph."

The Analogy: Imagine a library where books aren't just organized by title, but by how the pictures inside them relate to the words on the pages. If you ask for "evidence about rising temperatures," MEVER doesn't just look for the word "temperature"; it looks for books that mention heat AND books that contain line graphs showing upward trends. It connects the dots between what is said and what is shown, ensuring it finds the most relevant evidence.

2. The Master Detective (Claim Verification)

Once the evidence is on the desk, the AI has to compare the "Claim" (what the person said) against the "Evidence" (the text and the charts).

The Analogy: Think of this like a detective using a magnifying glass on two different things at once. Instead of looking at the text and then looking at the chart separately, MEVER performs "Token- and Evidence-level Fusion." It’s like a detective who can overlay a transparent sheet of the chart's data directly onto the written sentence. By "stacking" the visual data on top of the text, the AI can spot tiny discrepancies—like if a sentence says "sales doubled" but the bar chart only shows a tiny increase.

3. The Eloquent Lawyer (Explanation Generation)

A judge who says "Guilty!" without explaining why is a bad judge. MEVER uses a special module called "Fusion-in-Decoder" to write out its reasoning.

The Analogy: Imagine a lawyer who doesn't just shout a verdict, but carefully walks the jury through the evidence: "I am refuting this claim because while the text says X, the blue line in Figure 1 clearly shows Y." To make sure the lawyer doesn't accidentally contradict themselves, the researchers added a "Consistency Regularizer"—essentially a "sanity check" that ensures the written explanation actually matches the final verdict.

Why does this matter? (The AIChartClaim Dataset)

The researchers realized that most AI training is done on "general" stuff (like news or social media). But science is different! To test their AI, they created a brand-new "textbook" called AIChartClaim. It’s filled with complex scientific charts from AI research papers. This forces the AI to move beyond simple "fact-checking" and into the realm of high-level scientific reasoning.

The Bottom Line

MEVER is moving AI from being a simple "text-reader" to a "multi-sensory thinker." It doesn't just tell you if a claim is true or false; it finds the right proof, looks at the pictures, and explains its logic in a way that humans can actually trust.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →