← Latest papers
💻 computer science

FinDocMRE: A Benchmark for Document-Level Financial Multimodal Reasoning Evaluation

This paper introduces FinDocMRE, a comprehensive multi-image document-level benchmark comprising over 12,000 samples from financial reports to rigorously evaluate and expose the limitations of current Large Multimodal Models in integrating text, tables, and images for complex financial reasoning.

Original authors: Jiayong Zhu, Jiangtong Li, Jinru Ding, Dawei Cheng, Jie Xu, Feng Yu

Published 2026-05-19
📖 4 min read☕ Coffee break read

Original authors: Jiayong Zhu, Jiangtong Li, Jinru Ding, Dawei Cheng, Jie Xu, Feng Yu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new financial analyst. You have a massive, 100-page report filled with text, complex spreadsheets, and colorful charts scattered across different pages. You want to see if your new hire can look at the whole picture, connect the dots between a chart on page 5 and a table on page 40, and do the math correctly to give you a precise answer.

That is exactly what this paper, FinDocMRE, is about. It's a giant "final exam" designed to test how well Artificial Intelligence (AI) can handle these messy, real-world financial documents.

Here is a breakdown of what the researchers did and what they found, using simple analogies:

1. The Problem: The "Single-Chart" Trap

Until now, most AI tests were like showing a student a single, isolated math problem on a piece of paper. The AI could solve it easily. But in the real financial world, information isn't isolated. It's like a jigsaw puzzle where the pieces are scattered across a whole book.

  • The Issue: Existing AI models are great at looking at one chart and saying, "Oh, that line is going up." But they struggle when asked to say, "Take the number from the chart on page 10, multiply it by the percentage in the table on page 30, and tell me the total cost."
  • The Gap: There was no big, tough test that forced AI to do this "document-level" reasoning.

2. The Solution: Building the "FinDocMRE" Exam

The team created a massive new benchmark (a test set) called FinDocMRE. Think of it as building a gym for AI to train on heavy lifting.

  • The Dataset: They gathered 2,878 real financial reports (like annual reports from big companies) and pulled out 12,207 specific questions.
  • The "Recipe" for Quality: They didn't just ask a computer to write questions, because computers sometimes make up facts (a problem called "hallucination"). Instead, they used a three-step recipe:
    1. The Robot Chef: An AI generated the questions based only on the images and charts, ignoring the surrounding text to force it to "see" the data.
    2. The Human Taste-Testers: Three real financial experts (like senior analysts) reviewed every single question. If the question was confusing, the math was wrong, or the chart reference was off, they threw it in the trash.
    3. The Final Cut: Only about 55% of the generated questions made the cut. This ensured the final exam was incredibly high-quality and difficult.

3. The Results: The "Analyst vs. Calculator" Split

The researchers tested 11 of the smartest AI models available (including big names like GPT-5, Gemini, and Qwen) against this exam. They also brought in real human financial experts to see how the AI compared.

Here is what they discovered:

  • The Scoreboard: The best AI model only scored about 64 out of 100. The human experts scored about 79. There is still a huge gap.
  • The "Storyteller" vs. The "Math Whiz":
    • Good at Stories: The AI models are surprisingly good at writing long, coherent summaries. If you asked, "What is the general trend of this company?" they could write a great paragraph.
    • Bad at Math & Navigation: When asked to do precise calculations or find a specific number hidden in a chart on page 45 while looking at a chart on page 5, they often failed. They would get the numbers wrong or mix up which chart belonged to which year.
  • The "Lost in the Middle" Effect: As the documents got longer and the AI had to look at more images (like 3 or 4 different charts at once), the AI's performance dropped. It's like trying to remember three different phone numbers at once; the more you add, the more likely you are to mix them up.

4. The Big Takeaway

The paper concludes that while AI is getting better at "seeing" and "talking," it is still struggling to act like a professional financial analyst.

  • The Bottleneck: The main problem isn't that the AI can't read the words; it's that it can't reliably ground its reasoning in the specific visual evidence scattered across a complex document. It's like a student who knows the theory of accounting but keeps looking at the wrong line on the spreadsheet when doing the math.

In short: FinDocMRE is a rigorous stress test that shows current AI is a great "essay writer" but a shaky "calculator" when dealing with complex, multi-page financial reports. The paper suggests that for AI to truly replace human analysts, it needs to get much better at navigating these visual puzzles without getting lost.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →