← Latest papers
💬 NLP

When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents

This paper introduces "Multimodal Finance Eval," the first benchmark for French financial document understanding, revealing that while current vision-language models excel at text and table extraction, they struggle significantly with chart interpretation and suffer from error propagation in multi-turn conversational reasoning.

Original authors: Virginie Mouilleron, Théo Lasnier, Anna Mosolova, Djamé Seddah

Published 2026-03-17
📖 4 min read☕ Coffee break read

Original authors: Virginie Mouilleron, Théo Lasnier, Anna Mosolova, Djamé Seddah

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a team of very smart, super-fast robots. These robots are experts at reading books and looking at pictures. They can tell you what a story is about or describe a photo in seconds.

Now, imagine you hire these robots to be financial advisors for a French bank. Their job is to read thousands of pages of boring, complicated investment documents (like prospectuses) that are filled with legal text, giant spreadsheets, and colorful charts. They need to answer questions like, "How much does it cost to join this fund?" or "What happens if the market drops?"

The paper you shared is essentially a report card for these robots. The researchers built a special test called MULTIMODAL FINANCE EVAL to see if these robots are actually ready for the job, or if they are just bluffing.

Here is the breakdown of what they found, using some everyday analogies:

1. The Test: A "Financial Obstacle Course"

The researchers didn't just ask the robots simple questions. They created a massive obstacle course with 1,204 questions based on real French financial documents.

  • The Text: Long, dense legal paragraphs (like reading a 500-page rulebook).
  • The Tables: Complex grids of numbers (like a giant Excel sheet).
  • The Charts: Visual graphs showing trends (like a stock market line going up and down).
  • The Conversation: A back-and-forth chat where the robot has to remember what it said five minutes ago.

2. The Results: The "One-Trick Pony" Problem

The robots did surprisingly well in some areas, but failed miserably in others.

  • Reading Text (The "Bookworm" Mode): 🌟 Great!
    When the question was just about reading a sentence, the robots were like super-fast librarians. They got about 90% of the answers right. If you asked, "What is the entry fee?" they could find it instantly.
  • Reading Tables (The "Accountant" Mode): 📊 Pretty Good.
    When the data was in a neat grid, the robots did well (around 85%). They could look at a row and column and find the right number.
  • Reading Charts (The "Artist" Mode): 🎨 Terrible.
    This is where it got weird. When the information was in a graph or a picture, the robots got confused. Their accuracy dropped to between 34% and 62%.
    • The Analogy: Imagine showing a robot a picture of a mountain range and asking, "Which peak is the highest?" The robot might just guess because it's good at reading words, but it's bad at "seeing" the shape of the mountain. It struggles to understand that a line going up means "good" and a line going down means "bad."

3. The Big Disaster: The "Domino Effect"

The most shocking part of the paper is what happened when they made the robots have a conversation.

  • The Scenario: You ask the robot a question. It answers. Then you ask a follow-up question that depends on the first answer.
  • The Problem: If the robot makes one small mistake in the first answer, it gets stuck. It tries to build the next answer on top of that wrong fact.
  • The Result: The accuracy crashes to about 50% (basically a coin flip).
    • The Analogy: Imagine a game of "Telephone." If the first person whispers the wrong word, the whole chain is ruined. Even if you give the robot a bigger brain (more "parameters" or intelligence), it doesn't help. A super-smart robot is just as likely to get stuck in a loop of wrong answers as a smaller one. Once it trips, it can't get back up.

4. Why This Matters

The researchers found that while these AI models are amazing at finding information (like a search engine), they are very brittle when they have to think through a complex problem step-by-step, especially when pictures and conversations are involved.

  • The Good News: We can trust them to find a specific number in a document quickly.
  • The Bad News: We cannot trust them to act as a financial advisor in a live chat. If they make one small error, they might give you terrible financial advice, and they won't know they are wrong.

The Bottom Line

The paper concludes that we shouldn't just keep making these robots "bigger" (adding more brain power). Instead, we need to teach them how to double-check their work, how to understand pictures better, and how to recover when they make a mistake. Until then, using them for high-stakes financial advice is like letting a toddler drive a race car: they might be fast, but they aren't ready for the track.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →