← Latest papers
🤖 AI

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

This paper introduces ExtractBench, a comprehensive benchmark for evaluating schema-guided enterprise document extraction across 370 documents and 67 types, which reveals that LlamaExtract Agentic Plus achieves state-of-the-art accuracy and grounding at a significantly lower cost compared to coding agents.

Original authors: Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo

Published 2026-08-03
📖 3 min read☕ Coffee break read

Original authors: Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but instead of a crime scene, your crime scene is a messy pile of paperwork. Some papers are crisp and typed, others are crumpled receipts with scribbled handwriting, and some are so long they stretch across a whole bookshelf. In the world of artificial intelligence, there's a special kind of robot called an "agent" that is being taught to read these documents. Its job is to find specific clues—like a date, a dollar amount, or a name—and write them down in a neat, organized list called a "schema." Think of the schema as a strict recipe: the robot must follow it exactly, no matter how messy the ingredients (the documents) look.

But here's the tricky part: just because the robot writes down the right numbers doesn't mean it actually saw them. It might have guessed, or it might have missed a whole page of clues. For businesses, this is a big deal. If a robot is supposed to read thousands of insurance claims or bank statements, it needs to be not only accurate but also able to point its finger at the exact spot on the page where it found the answer. If it gets it wrong, a human needs to be able to quickly check the evidence. This is the challenge of "schema-guided extraction": getting a robot to follow a recipe perfectly, even when the ingredients are messy, long, or written in crayon.

Enter ExtractBench, a new, super-tough test designed to see which AI robots are truly ready for the real world. The researchers behind this study didn't just want to know if the robots could read; they wanted to know if they could read everything correctly, show their work, and do it without costing a fortune. They built a massive "exam hall" containing 370 different types of documents—ranging from short receipts to 50-page government reports—covering 8 different industries like finance, energy, and healthcare. They even included documents that were scanned, handwritten, or had weird tables that spanned multiple pages.

The big surprise? The "smartest" looking robots didn't always win. The study found that while some powerful AI models are great at reading short, clean documents, they often get tired and stop reading halfway through long ones, missing huge chunks of data. On the other hand, some robots that write their own code to solve the problem are very accurate but cost a lot of money to run. The researchers discovered a "sweet spot" with a system called LlamaExtract Agentic Plus. It managed to get the right answers almost as often as the expensive code-writing robots, but it did it for a fraction of the price.

However, the study also uncovered a major blind spot. Even the best robots struggle with "grounding"—the ability to point to the exact word on the page that gave them the answer. While some systems could tell you which page the answer was on, fewer than half could point to the exact word. The paper suggests that while we are getting better at getting the right numbers, teaching robots to reliably show their homework is still a work in progress. In short, ExtractBench proves that for enterprise work, you don't just need a smart robot; you need one that is thorough, honest about where it found its clues, and doesn't break the bank.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →