GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents
The paper introduces GDP.pdf, a new benchmark comprising 100 professionally authored question-document pairs across ten fields that reveals significant limitations in current frontier multimodal models, with the best-performing model achieving only a 15% pass rate due to recurring errors in table/chart interpretation, cross-referencing, and handling document amendments.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're playing a high-stakes game of "Find the Clue" inside a massive, messy library. But instead of books, the library is filled with PDFs: insurance policies, construction blueprints, medical guidelines, and employee handbooks. You have a team of super-smart robot detectives (the AI models) ready to solve the mystery. You'd think, "With all their power, they'll ace this!"
But here's the twist: when the researchers at Surge AI put these robots to the test, the results were a total shocker.
The Big Reveal: The Robots Are Getting Lost
The paper introduces a new challenge called GDP.pdf. Think of it as a "final boss" level for AI, designed specifically to see if robots can actually do the boring, tricky paperwork that real humans do every day. The researchers gathered 100 real-world documents from 10 different jobs (like finance, healthcare, and construction) and asked the robots questions that a real human would ask.
The result? Even the smartest robot detectives in the world failed miserably. The very best model only got 15% of the questions right. The worst one? It only passed 1% of them. That means if you asked the best robot to read a 10-page insurance policy and find a specific rule, it would get the answer wrong 85 times out of 100.
Why Did They Fail? (The "Gotchas")
The paper argues that previous tests were like giving the robots a clean, easy-to-read textbook. But real PDFs are messy. They are full of traps that the robots couldn't handle. Here are the specific ways they tripped up:
- The "Footnote" Trap: Imagine a rule is written in the main text, but then a tiny footnote three pages later says, "Actually, ignore that rule for this specific case." The robots read the main text, gave a confident answer, and completely missed the footnote. It's like reading a map, seeing a road, and ignoring the small sign that says "Road Closed."
- The "Amendment" Mix-up: Sometimes, a document has a new page attached that changes an old rule. The robots would quote the old rule as if it were still true, even though the new page had officially erased it.
- The "Table" Tangle: When numbers are in a grid with merged cells or lines that cross pages, the robots got confused. They'd look at the right row but the wrong column, or mix up the headers. It's like trying to read a menu where the prices are floating in the air instead of next to the food.
- The "I Know Better" Problem: This is the funniest (and most dangerous) failure. If a robot's training data told it "Drill bits work this way," but the specific PDF said "Do NOT use those drill bits for this metal," the robot would ignore the PDF and give the answer it "knew" from its training. It trusted its memory over the actual document in front of it.
What the Paper Says We Can't Do Yet
The authors are very clear about what this means. They explicitly rule out the idea that these models are ready to work alone on professional documents. Just because a robot scores high on a standard, academic test (like identifying a cat in a picture or answering a simple math question) doesn't mean it can handle a real-world lease agreement.
The paper suggests that simply making the robots' "memory" (context windows) bigger won't fix this. The problem isn't that they can't see the whole document; it's that they can't understand the messy structure, the tiny footnotes, and the visual clues like charts and floor plans.
How Sure Are We?
The researchers didn't just guess; they measured this. They built a strict scoring system where a robot has to get every single part of the answer right to pass. They tested seven of the most advanced models available as of April 2026. The results were consistent: the robots are currently not reliable enough to be trusted with unsupervised work on these documents.
The Takeaway
Think of GDP.pdf as a reality check. It's a benchmark that says, "Hey, before we let AI handle our legal contracts or medical records, we need to teach it how to read the fine print, not just the big headlines." Until the robots can stop ignoring those tiny footnotes and stop guessing when the document says otherwise, they aren't ready for the big leagues of professional paperwork. The gap between what these models can do in a classroom and what they can do in the real world is still huge.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.