← Latest papers
💻 computer science

Evaluating Vision-Language Models for Structured Information Extraction from Heterogeneous Multi-Page Laboratory Reports

This study evaluates vision-language models for extracting structured data from heterogeneous multi-page laboratory reports, finding that while whole-document processing offers superior speed, a page-by-page workflow using Qwen2.5-VL achieves the highest accuracy and stability across varying document lengths.

Original authors: Ashish Katyal, Ashwani Bhatnagar

Published 2026-09-25
📖 5 min read🧠 Deep dive

Original authors: Ashish Katyal, Ashwani Bhatnagar

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Hospitals generate a vast ocean of paper and digital records every day, from the narrative notes doctors write to the dense, tabular results of laboratory tests. For computers to help doctors make better decisions or to help researchers find patterns in disease, this information must be converted from images and text into clean, organized data that machines can read. This process, known as information extraction, has long been a stumbling block because medical documents are messy. They come in different shapes, sizes, and formats; some are crisp digital files, while others are grainy scans of old paper. Furthermore, the information is often scattered across multiple pages, with test names, numbers, and units arranged in complex tables that look different from one hospital to the next.

In recent years, a new type of artificial intelligence called a vision-language model has emerged to tackle this challenge. Unlike older systems that could only read text or only see images, these models can look at a picture of a document and understand both the visual layout and the words within it simultaneously. They can see that a number belongs to a specific test because of where it sits on the page, not just because it follows a specific word. However, while these models are powerful, scientists did not fully understand the best way to feed them these complex, multi-page documents. Should the computer look at an entire ten-page report at once, or should it examine the pages one by one? The answer to this question determines whether the computer misses important details or wastes time, a distinction that matters deeply when the data is used to guide patient care.

A team of researchers set out to find the answer by running a controlled experiment with real-world laboratory reports. They gathered de-identified test results from two major providers, Lal PathLabs and Thyrocare, and created versions of these reports in both digital and scanned formats. To ensure a fair test, they prepared three different lengths of documents: a short single-page report, a medium report of five to six pages, and a long report stretching to ten pages. They then tested three different ways of processing these documents using two leading artificial intelligence models, Qwen2.5-VL and Llama 3.2 Vision. The first approach asked the model to look at the entire report as a single image. The second approach asked the model to look at each page individually and then combine the results. The third approach used a different model to look at pages individually.

The results revealed that the strategy used to present the document was just as important as the model itself. When the researchers asked the Qwen2.5-VL model to process the reports page by page, it achieved the highest level of accuracy, correctly identifying and recording nearly 99 percent of the expected test results. This method proved particularly robust for the longest documents; even with ten pages of data, the page-by-page approach maintained an accuracy of over 98 percent. In contrast, when the same model was asked to view the entire ten-page report at once, its accuracy dropped significantly to about 91 percent. The model tended to miss tests that appeared on later pages when the document was too long to view all at once.

The trade-off for this higher accuracy was time. The page-by-page method took roughly twice as long to process a report as the method that viewed the whole document at once. While the complete-document approach was faster and extremely precise when it did find a test, it simply failed to see many of the tests hidden in longer reports. The third workflow, which used the Llama 3.2 Vision model to look at pages one by one, performed the worst of all. It took the longest to finish, averaging over 108 seconds per report, and it frequently generated incorrect records or missed data, achieving an overall accuracy of less than 88 percent. This suggested that simply breaking a document into pieces was not enough; the specific intelligence of the model mattered just as much as the method of presentation.

The study also examined how the quality of the document affected the results. Whether the report was a clean digital file or a scanned image of a paper document made little difference to the performance of the best-performing workflow. The page-by-page approach with the Qwen2.5-VL model maintained perfect accuracy under the evaluated scanned conditions, which included short and medium-length reports. This finding suggests that the physical quality of the scan is less critical than the strategy used to feed the information to the computer. The researchers noted that the specific layout of the hospital's report also played a role, with one provider's reports being slightly easier to process than the other's, but the overall trend held true across both sources.

Ultimately, the research demonstrates that there is no single "best" way to extract data from medical reports. The choice depends entirely on what the user needs. If a hospital system needs to process thousands of reports quickly and can tolerate missing a few details that can be caught later, the faster, whole-document approach might be suitable. However, if the goal is to build a complete and reliable record of every single test result, especially from long and complex documents, the slower, page-by-page method is the superior choice. The study concludes that the way a document is presented to an artificial intelligence system is a critical design decision that directly impacts the reliability of the medical data it produces.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →