← Latest papers
💬 NLP

Document-as-Image Representations Fall Short for Scientific Retrieval

This paper argues that document-as-image representations are suboptimal for scientific retrieval and introduces the ArXivDoc benchmark to demonstrate that leveraging structured LaTeX sources, particularly text-based representations, significantly outperforms image-based approaches for retrieving evidence from text-rich scientific documents.

Original authors: Ghazal Khalighinejad, Raghuveer Thirukovalluru, Alexander H. Oh, Bhuwan Dhingra

Published 2026-04-21
📖 4 min read☕ Coffee break read

Original authors: Ghazal Khalighinejad, Raghuveer Thirukovalluru, Alexander H. Oh, Bhuwan Dhingra

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to find a specific recipe in a massive library of cookbooks.

The Old Way (Document-as-Image):
For a long time, researchers thought the best way to search these libraries was to take a photo of every single page, feed those photos into a computer, and ask it to find the right one. It's like handing a librarian a blurry snapshot of a page and saying, "Find me the book with this picture."

The problem? If the recipe you need is hidden in a tiny footnote, a complex chart, or a paragraph of text, the computer just sees a blob of pixels. It has to guess what the words say and where the ingredients are listed. As the book gets thicker and the pages get more crowded with text, this "photo search" gets worse and worse. It's like trying to read a novel by squinting at a photograph of the cover.

The New Way (ArXivDoc):
The authors of this paper, Document-as-Image Representations Fall Short for Scientific Retrieval, decided to test a different approach. They realized that scientific papers (like those on ArXiv) are actually built from a digital blueprint called LaTeX. Think of LaTeX not as a picture, but as the raw ingredients list and the chef's notes before the dish is plated.

Instead of taking a photo of the finished plate, they looked at the blueprint. This allowed them to see exactly where the text, the math equations, the tables, and the figures were located, without any guesswork.

The Big Experiment:
The team built a new "library" called ArXivDoc containing over 8,000 scientific papers. They asked a simple question: Is it better to search using the raw text/blueprint, or the photos of the pages?

They tested three main strategies:

  1. The Photo Search: Using the computer vision models to scan page images.
  2. The Text Search: Using the raw text (and adding AI descriptions of the pictures).
  3. The Mixed Search: A combination of text and pictures, kept in the order they appear in the story.

The Surprising Results:

  • Photos Lost: The "Photo Search" was consistently the worst performer. Even when people were looking for a specific picture or chart, the computer did better if it just read the text around the picture. Why? Because scientists usually describe their charts in detail right next to them. The text holds the secret; the picture is just the illustration.
  • Text is King: The best results came from reading the text. Even for questions about diagrams, the text description was enough to find the right paper. It's like finding a book by reading the summary on the back cover rather than trying to recognize the font on the title page.
  • The "Mosaic" Approach: The most effective method was a mix of text and images, but not as a flat photo. Imagine a mosaic where you keep the tiles (text and images) separate but arranged in the correct order. This allowed the computer to understand the story flow without getting confused by the visual clutter of a full page.

The "Long Book" Problem:
The researchers also found that as books get longer (more pages, more text), the "Photo Search" crashes. It's like trying to find a specific sentence in a 500-page novel by looking at a single photo of the whole book; the details get lost. The text-based search, however, stays sharp no matter how long the book is.

The Takeaway:
The paper argues that for scientific documents, we shouldn't treat them like art galleries (images) but like libraries (text).

Current AI models are trained to look at pages as pictures, but scientific papers are actually structured data. By switching to the "blueprint" (text and structured elements) and treating images as supporting characters rather than the main event, we can find the right scientific answers much faster and more accurately.

In a nutshell: Don't just look at the picture of the map; read the legend and the coordinates. The text holds the truth, and the pictures are just there to help you visualize it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →