MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
This paper introduces MDPBench, the first benchmark for multilingual document parsing across 17 languages and diverse real-world conditions, revealing that while closed-source models remain robust, open-source alternatives suffer significant performance drops on photographed documents and non-Latin scripts.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a massive library containing books, receipts, and notes written in 17 different languages. Some of these are perfect, crisp digital files you can copy and paste. Others are messy: they are photos of crumpled papers taken in dim light, pages bent by a heavy book, or receipts snapped on a bumpy bus ride.
For a long time, the "librarians" (AI models) trying to read these documents have been like students who only studied in a perfect, quiet classroom. They are great at reading clean, digital English or Chinese text. But if you handed them a photo of a crumpled Arabic receipt or a Thai menu taken in the rain, they would often panic, guess wrong, or just give up.
Enter MDPBench: The "Real-World Stress Test" for AI.
This paper introduces a new tool called MDPBench (Multilingual Document Parsing Benchmark). Think of it as a giant, rigorous driving test for AI models, but instead of driving a car, they are trying to "read" documents.
Here is the breakdown of what they did and what they found, using some simple analogies:
1. The Test Track (The Dataset)
The researchers built a test track with 3,400 document images.
- The Languages: They didn't just stick to the usual suspects (English, Chinese, Spanish). They included 17 languages, including tricky ones like Arabic, Hindi, Thai, and Russian.
- The Conditions: They took perfect digital documents and deliberately "ruined" them to simulate real life. They printed them out, crumpled them, bent them, took photos of them in the dark, in the sun, and with weird camera angles.
- The Goal: To see if AI can read a document when it's not perfect.
2. The Judges (The Annotation)
You can't grade a test unless you have the answer key. The researchers didn't just ask one AI to write the answers (because AIs make mistakes). Instead, they used a "Three-Layer Safety Net":
- The Experts: Three different top-tier AI models tried to read the document.
- The Consensus: If two out of three agreed on a word, that was likely correct. If they all disagreed, they used a super-smart AI (Gemini-3-Pro) to double-check.
- The Humans: Real human experts then reviewed the work, fixing any remaining errors. This ensures the "answer key" is perfect.
3. The Results: Who Passed the Test?
The researchers put various AI models through this stress test, from free, open-source models to expensive, closed-source ones.
- The Star Performer: The Gemini-3-Pro (a closed-source, paid model) was the only one that really handled the chaos. It was like a seasoned driver who could navigate a pothole-filled road without losing control. It got about 86% of the answers right.
- The Struggling Students: The open-source models (the free ones available to everyone) did much worse. They were like students who only practiced on smooth, dry roads. When faced with a crumpled photo or a non-Latin script (like Arabic or Hindi), their performance crashed.
- On photographed documents, their scores dropped by nearly 18%.
- On non-Latin scripts, they dropped by 14%.
4. The Specific Glitches (Where They Failed)
The paper highlights some funny but frustrating ways these AIs fail in the real world:
- The "Right-to-Left" Confusion: Arabic is read from right to left. Many models tried to read it like English (left to right), resulting in gibberish. It's like trying to read a mirror image without a mirror.
- The "Space" Hallucination: In Thai, words aren't separated by spaces. The AI kept inventing spaces where they didn't belong, breaking words into nonsense pieces (like reading "biggest" as "big ge st").
- The "Copy-Paste" Loop: Some models got stuck in a loop, repeating the same sentence over and over again, or hallucinating text that wasn't there at all.
- The "Language Drift": Sometimes, the AI would look at a Vietnamese document and confidently say, "This is Chinese!" because it was guessing based on patterns it saw in its training data.
Why Does This Matter?
Currently, most AI is built for the "clean digital world." But the real world is messy. Most of the world's knowledge isn't in perfect PDFs; it's in old books, handwritten notes, and photos of receipts.
MDPBench is a wake-up call. It shows us that while our AI is getting smarter, it's still fragile. It works great in the lab but struggles in the field. This benchmark gives researchers a clear map of where the AI is failing so they can build systems that are truly ready for the messy, multilingual, real world.
In short: We built a test to see if AI can read a messy, crumpled receipt in 17 different languages. The expensive AI passed with flying colors; the free AI mostly failed. Now we know exactly what to fix.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.