MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc
MMLongBench-Doc-V2 is a corrected and semantics-aware revision of the MMLongBench-Doc benchmark that fixes 106 ground-truth annotations, replaces string-matching metrics with an LLM-based judge, and removes invalid samples to provide a more reliable evaluation framework for long-document QA.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a judge in a massive, high-stakes talent show where the contestants are super-smart computer programs called "AI." These AIs are trying to prove they can read thick, boring, and complicated documents—like financial reports, scientific papers, and instruction manuals—and answer questions about them. This is the world of Document QA (Question Answering). In this arena, the goal is simple: the AI reads a long PDF, and you ask it a question. If it gets the answer right, it wins points. If it gets it wrong, or if it makes up an answer when the document doesn't have the information, it loses points.
Why do we care? Because as these AIs get smarter, they start acting like overconfident students who guess the answer just to look busy. We need a way to test if they are actually reading or just hallucinating (making things up). To do this, scientists created a giant test called MMLongBench-Doc. It's like a final exam with over 1,000 questions spread across 135 different documents. But, as this new paper reveals, the exam itself had some serious flaws that made the grades unfair.
The Flawed Exam: When the Grader is the Problem
Think of the original exam (MMLongBench-Doc) as a strict, old-fashioned teacher who grades with a red pen that only looks for exact spelling matches. If the correct answer is "1,358,000" and the AI writes "1358000" (without the commas), the teacher marks it wrong. If the answer is "Operating Activities" and the AI writes "Operations activities," the teacher marks it wrong. It's like a math teacher failing a student for writing "12" instead of "12.0," even though the number is right.
Even worse, the teacher's answer key was sometimes wrong. Imagine a question asking, "How many blue arrows are on page 5?" The answer key says "Zero." But if you look at the page, there are actually no blue arrows, so "Zero" is correct. However, sometimes the key said "Zero" when the document actually had an arrow, or the question was written so confusingly that even a human couldn't answer it. The worst part? These mistakes were concentrated on the hard questions. So, the smartest AIs were getting penalized the most, while the dumber ones (who guessed randomly) weren't affected as much. It was like a race where the fastest runners were weighed down by lead boots, while the slow walkers weren't.
The Fix: MMLongBench-Doc-V2
Enter the new paper: MMLongBench-Doc-V2. The author, who is Mingtian Zhang, decided to fix the exam rather than throw it away. They treated the original test like a messy draft and cleaned it up with three major moves.
1. The New Grader: A "Meaning" Detective
Instead of a robot that just checks if two strings of text look identical, they hired a "Meaning Detective" (a special AI judge). This detective doesn't care about commas, capital letters, or extra spaces. It looks at the AI's answer and asks, "Does this mean the same thing as the correct answer?"
- The Analogy: If the answer is "The cat is sleeping," and the AI says "A feline is napping," the old grader would fail it. The new detective gives it a pass because the meaning is the same.
- The Catch: This detective never sees the original document. It only compares the AI's answer to the correct answer key. This prevents the detective from getting confused by bad scans or missing text in the PDF.
2. Fixing the Answer Key (106 Corrections)
The author went through the exam and found 106 questions where the answer key was wrong, the question was broken, or the document didn't match the question.
- The "Wrong File" Problem: They found 10 questions that were asking about a document that wasn't even in the test folder! It was like asking a question about "Chapter 5 of Harry Potter" but handing the student a copy of The Hobbit. No one could answer that, so they removed those questions entirely.
- The "Sign" Error: In one case, the key said a number went up by 60.3%, but the document showed it went down. The author fixed the sign.
- The "Ambiguous" Fix: Some questions were too vague. If a document listed "short-term debt" and "long-term debt" separately, but the question asked for "total debt" without saying which ones to add, the author rewrote the question to be super specific.
3. The "Empty Set" Dilemma
This is the trickiest part. Some questions ask, "How many dragons are in this financial report?" The answer is obviously zero. But in the original test, some "zero" answers were marked as "Not Answerable" (meaning the document didn't have the info), while others were marked as "0" (meaning the document was checked and found nothing).
- The Rule: The author created a strict rule: If the document exists and you can prove the thing isn't there, the answer is 0. If the document is missing a whole page or the question asks about something that can't be counted (like "all the stars in the universe"), then it's Not Answerable.
- The Result: They adjusted 14 questions to make sure the "zero" answers were fair. This stops the AI from getting tricked into thinking "I don't know" is the right answer when the right answer is actually "None."
The Final Scoreboard
After all the cleaning, the new test (V2) has 1,071 questions across 134 documents (down from 1,082 questions and 135 documents).
- The Big Change: The scores on this new test are not comparable to the old scores. You can't say, "The AI got 44.9% last year and 50% this year, so it improved." The test itself changed, so the numbers are on a different scale.
- The Good News: The new test is a fairer way to see if an AI is actually smart or just lucky. It removes the "trick questions" that were confusing the best models.
- The Caveat: The author admits they didn't check every single question. They focused on the ones that smart AIs got wrong, which is where the errors were most likely to hide. There might still be a few mistakes left in the test, but they are much fewer than before.
Why This Matters
This paper isn't about inventing a new super-AI. It's about fixing the ruler we use to measure them. If you are building a robot to read legal contracts or medical reports, you need to know if it's actually reading or just guessing. MMLongBench-Doc-V2 gives us a cleaner, more honest ruler. It ensures that when an AI gets a question right, it's because it understood the document, not because it guessed the right number of commas.
The author has made all their corrections, the new test data, and the tools to grade the answers available for everyone to use. It's a reminder that in science, sometimes the most important work isn't building the engine, but fixing the speedometer so we know how fast we're really going.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.