← Latest papers
💻 computer science

How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings

This paper introduces PureDocBench, a source-traceable benchmark covering clean, degraded, and real-world document images that exposes significant flaws in the saturated OmniDocBench rankings and reveals that document parsing remains an unsolved challenge with substantial performance gaps between models and persistent bottlenecks in formula recognition.

Original authors: Zhiheng Li, Zongyang Ma, Jiaxian Chen, Jianing Zhang, Zhaolong Su, Yutong Zhang, Zhiyin Yu, Ruiqi Liu, Xiaolei Lv, Bo Li, Jun Gao, Ziqi Zhang, Chunfeng Yuan, Bing Li, Weiming Hu

Published 2026-05-12
📖 4 min read☕ Coffee break read

Original authors: Zhiheng Li, Zongyang Ma, Jiaxian Chen, Jianing Zhang, Zhaolong Su, Yutong Zhang, Zhiyin Yu, Ruiqi Liu, Xiaolei Lv, Bo Li, Jun Gao, Ziqi Zhang, Chunfeng Yuan, Bing Li, Weiming Hu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to judge who is the best chef in the world. For the past year, everyone has been competing in a single, tiny kitchen called OmniDocBench. This kitchen has only 1,355 dishes, and the judges (the scoring system) have been so generous that almost every top chef gets a 90% or higher. It's like a race where everyone is crossing the finish line at the exact same time; you can't tell who is actually faster or better.

Furthermore, the judges made some mistakes. They miscounted ingredients, missed steps, and even hallucinated flavors that weren't there. Because the "recipe" (the ground truth) was flawed, the leaderboard rankings are unreliable. Plus, since everyone has been cooking in this same kitchen for a year, the chefs might have memorized the dishes rather than learning how to cook.

Enter PureDocBench, the new, massive, and perfectly organized kitchen introduced in this paper.

The New Kitchen: PureDocBench

Instead of using old, messy recipes, the creators of PureDocBench built their kitchen from scratch using digital blueprints (HTML/CSS).

  • The Blueprint: They wrote the code for the document first.
  • The Dish: They used that code to generate the image of the document.
  • The Answer Key: Because they wrote the code first, they know exactly what the text, tables, and formulas should look like. There is no guessing.

This kitchen is huge. It has 10 different types of restaurants (Academic, Legal, Medical, Finance, etc.) and 66 specific menu items (like invoices, diplomas, or lab reports).

The Three Versions of Every Dish

To test how tough the chefs really are, the kitchen serves every single dish in three different conditions:

  1. Clean: A perfect, fresh plate straight from the oven.
  2. Digitally Degraded: The plate was scanned by a bad scanner, or the photo was compressed, making it look blurry or yellowed (like an old fax).
  3. Real-World Degraded: The plate was printed out, then someone took a photo of it with a shaky hand in dim lighting, or maybe it was a photocopy of a photocopy.

This is crucial because a chef might be great at plating a perfect dish but terrible at serving a messy one.

The Results: Who Actually Won?

The researchers tested 40 different chefs (AI models) in this new kitchen. Here is what they found:

1. The Race is Far From Over
In the old kitchen, the best chefs scored over 90%. In PureDocBench, the absolute best chef only scored 74 out of 100. There is a massive gap (44 points) between the best and the worst. This means document parsing is not solved yet; there is still a lot of room for improvement.

2. Small Specialists vs. Giant Generalists
There are two types of chefs:

  • Specialists: Chefs who only know how to cook documents. They are small and efficient.
  • Generalists: Giant, multi-talented chefs who can cook anything but aren't specialized in documents.

The paper found that the small specialists (some with less than 4 billion "brain cells") are just as good as, or even better than, the giant generalists (which can be 100 times larger). It's like a small, specialized sushi chef beating a massive, all-you-can-eat buffet chef at making sushi.

3. The Formula Bottleneck
No matter how smart the chef is, math formulas are the hardest part. Even the best models only get about 67% right on formulas. It's a universal weak spot.

4. The "Messy Plate" Surprise
This is the most interesting finding.

  • The Specialist chefs crumbled when the plates got messy (Real-World Degraded). Their scores dropped significantly.
  • The Generalist chefs were surprisingly tough. They handled the messy, blurry, real-world photos much better.

This means that if you only test chefs on perfect, clean plates, you get the wrong idea about who will actually succeed in the real world. The rankings flip when you add real-world noise!

The Takeaway

The paper argues that we need to stop using the old, flawed kitchen (OmniDocBench) to judge AI. We need to use PureDocBench, which:

  • Has a perfect, traceable answer key (no more judge errors).
  • Tests chefs on messy, real-world conditions, not just perfect ones.
  • Reveals that while AI is getting better, it still has a long way to go, especially with math formulas and handling real-world messiness.

The researchers have opened the doors to this new kitchen, giving everyone the blueprints, the dishes, and the answer keys so the community can build better chefs together.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →