← Latest papers
💻 computer science

EvidenceLens: A Claim-Evidence Matrix for Auditing Financial Question Answering

EvidenceLens is a visual analytics prototype that enhances the auditability of financial question answering by decomposing model outputs into atomic claims and visualizing their alignment with multimodal source evidence through a structured claim-evidence matrix.

Original authors: Fengchen Gu, Xiaotian Ren, Zhengyong Jiang, Zhilu Zhang, Ángel F. García-Fernández, Angelos Stefanidis, Mian Zhou, Huakang Li, Jionglong Su

Published 2026-06-24
📖 4 min read☕ Coffee break read

Original authors: Fengchen Gu, Xiaotian Ren, Zhengyong Jiang, Zhilu Zhang, Ángel F. García-Fernández, Angelos Stefanidis, Mian Zhou, Huakang Li, Jionglong Su

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a financial analyst, and you ask a super-smart AI assistant: "Did the company's profits go up because they got better at making things, or just because they sold off an old building?"

The AI gives you a confident, polished paragraph saying, "Yes, it was definitely because they got better at making things." It sounds perfect. But in the high-stakes world of finance, "sounds perfect" isn't good enough. You need to know if the AI is telling the truth or just making things up to sound smart.

This is where EVIDENCELENS comes in. Think of it as a "Truth Detector" or a "Fact-Checking X-Ray" for AI answers.

The Problem: The "Smooth Lie"

Current AI chatbots are like smooth-talking salespeople. They can blend real facts, weak guesses, and total inventions into one seamless story. If you just read the answer, it looks great. But if you look closely, you might find that the AI is ignoring a chart that says the opposite, or it's ignoring a tiny footnote that ruins the whole story.

The Solution: Breaking the Answer into Lego Bricks

EVIDENCELENS changes the game by refusing to treat the AI's answer as one big block of text. Instead, it breaks the answer down into tiny, atomic "Claims" (like individual Lego bricks).

  • Claim 1: "Profits went up."
  • Claim 2: "This was because of better manufacturing."
  • Claim 3: "Selling the building didn't matter."

Once the answer is broken down, the system checks every single brick against the original documents (the annual report, the spreadsheets, and the charts).

The Magic Tool: The "Claim-Evidence Matrix"

The heart of EVIDENCELENS is a big grid (a matrix) that looks a bit like a spreadsheet or a game board.

  • The Rows are the AI's claims.
  • The Columns are the evidence found in the documents (text, tables, or charts).

This grid uses colors to tell a story instantly:

  • Green: "Great! This claim is backed up by a table and a chart."
  • Yellow/Orange: "Hmm. This claim is supported by text, but there's no number to back it up."
  • Red: "Stop! The AI says one thing, but the chart shows the exact opposite."
  • Empty Space: "The AI made this up. There is no evidence for this claim at all."

How It Helps You (The Analyst)

Imagine you are a detective. Instead of reading a 50-page report to find one lie, EVIDENCELENS hands you a map that highlights exactly where the trouble is.

  1. Spotting the "Fake Confidence": Sometimes the AI sounds very sure, but the evidence is weak. EVIDENCELENS highlights this "Confidence Gap." It's like a teacher grading a student who writes a long, fancy essay but gets the math wrong. The system flags it immediately.
  2. Finding the "Hidden Contradictions": Maybe the main text says one thing, but a tiny footnote or a specific bar on a chart says another. The matrix shows these conflicting colors side-by-side, so you don't miss the contradiction hidden in the fine print.
  3. Checking the "Ingredients": It shows you if the AI is relying only on text (which can be vague) or if it actually looked at the hard numbers in the tables and the trends in the charts.

The Real-World Test

The paper tested this with two real scenarios:

  1. The "Profit Boost" Case: The AI confidently said profits rose due to efficiency. EVIDENCELENS showed that while the fact of rising profits was true, the reason (efficiency) was only half-supported, and the AI completely ignored evidence that a one-time sale of an asset was actually the main cause.
  2. The "Guidance Cut" Case: The AI said a drop in orders meant customers were leaving forever. EVIDENCELENS showed that while the text said one thing, the chart only showed a temporary dip, not a permanent collapse.

The Bottom Line

EVIDENCELENS doesn't just give you an answer; it gives you the audit trail. It turns a "black box" AI that spits out text into a transparent system where you can see exactly which parts of the answer are solid, which are shaky, and which are made up.

In the paper's own words, it treats financial questions not as a "writing problem" (how to write a good sentence) but as an "alignment problem" (does the sentence actually match the evidence?). It's a tool to help humans stay in the driver's seat when using AI for important financial decisions.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →