ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution
This paper introduces ViTaB-A, an evaluation framework demonstrating that while Multimodal Large Language Models can answer questions about structured data, they currently lack the reliability to provide accurate, fine-grained attribution for their answers across various formats, thereby limiting their trustworthiness in transparency-critical applications.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, well-read assistant who can look at a complex spreadsheet (like a company's financial report or a medical chart) and answer your questions about it.
You ask, "Did the company's revenue go up in 2023?"
The assistant says, "Yes, it did!"
That sounds great, right? But here is the catch: Where did they get that answer?
In the real world, especially in fields like law, medicine, or finance, you can't just take an answer on faith. You need to see the proof. You need to point to the exact row and column in the spreadsheet that proves the revenue went up.
This paper, titled ViTaB-A, is a report card on how good today's AI assistants are at doing exactly that: pointing to the proof.
The Big Discovery: The "Smart but Clueless" Assistant
The researchers tested several top-tier AI models (the "brains" behind these assistants) using three different ways to show them data:
- Images: A picture of a table.
- Markdown: A text file that looks like a table.
- JSON: A raw code format that computers love but humans find hard to read.
Here is what they found, using a simple analogy:
The "Chef" Analogy
Imagine you ask a chef, "Is the soup salty?"
- The Answer: The chef tastes it and says, "Yes, it is." (This is Question Answering).
- The Proof: You ask, "Show me the salt shaker you used." The chef points to a pepper shaker, or a random spoon, or just shrugs. (This is Attribution).
The study found that while the chefs (AI models) are getting the "Yes/No" answers mostly right (about 50-60% accuracy), they are terrible at pointing to the specific ingredient (the cell in the table) that led to that answer.
- For Images: The chefs are okay at pointing to the proof (about 33% accuracy).
- For Text (Markdown): They get worse.
- For Code (JSON): They are basically guessing. Their accuracy drops to near 0%. It's like asking a human to find a specific word in a book where every page is written in invisible ink.
The "Row vs. Column" Mystery
The researchers also noticed a funny pattern. The AI models are much better at finding the Row (the horizontal line, like a specific person's record) than the Column (the vertical category, like "Salary" or "Date").
Think of it like a classroom:
- Rows are the students. "Find John." (Easy! The AI can see the name).
- Columns are the subjects. "Find the Math score." (Hard! The AI often gets confused about which vertical line represents Math).
The AI is great at saying, "Here is John's record," but often fails to say, "And here is the specific box where his Math score lives."
The "Confidence Trap"
Here is the scariest part. The researchers asked the AI: "How sure are you that you found the right proof?"
The AI would say, "I am 90% sure!"
But when the researchers checked, the AI was actually wrong.
It's like a student taking a test who is loudly confident about their answer, even though they picked the wrong option. The study found that the AI's "confidence meter" is broken when it comes to proving its work. You cannot trust the AI's self-assessment to tell you if it's actually telling the truth about where it found the data.
Why Does This Matter?
If you are asking a chatbot to summarize a movie plot, it's fine if it gets the details slightly wrong. But if you are a doctor asking an AI to check a patient's drug dosage, or a judge asking it to review a legal contract, you need to know exactly where the AI got its information.
If the AI says, "The patient is allergic to Penicillin," but it can't point to the specific line in the medical record that says so, you cannot trust that answer.
The Takeaway
The paper concludes that current AI models are like magicians who can pull a rabbit out of a hat (give the right answer) but have no idea how they did it (can't show the trick).
They are getting better at guessing the right answer, but they are still very bad at showing their homework. Until we fix this, we have to be very careful using these tools for important, real-world decisions where "showing your work" is just as important as getting the answer right.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.