Scalable Visual Pretraining for Language Intelligence
This paper challenges the text-only paradigm of foundation model pretraining by demonstrating that unsupervised visual pretraining on raw visual documents, which preserves rich information like figures and layouts, consistently outperforms text-only approaches and offers a scalable pathway to enhanced language intelligence.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're trying to teach a super-smart robot how to understand the world. For a long time, the standard recipe has been to feed it mountains of books, articles, and websites, but with a catch: before the robot reads them, humans (or computers) strip away all the pictures, the squiggly math symbols, the charts, and the fancy page layouts, turning everything into plain, boring text. It's like trying to teach someone how to bake a cake by only giving them a list of ingredients written on a napkin, while throwing away the photo of the finished cake, the diagram of the mixing bowl, and the colorful graph showing how the oven temperature changes.
A new paper from 2026 suggests this "text-only" approach is missing a huge chunk of the puzzle. The researchers, led by a team from Shanghai and Zhejiang, argue that by throwing away the visual parts of documents, we are throwing away the very clues the robot needs to become truly smart. They call their new method Visual Pretraining (VP), and it's like teaching the robot to read the actual pages of a textbook, complete with diagrams and equations, rather than just the text description of them.
The Big Discovery: Seeing is Believing (and Learning)
The team tested this idea on some of the smartest AI models available (like Qwen and Llama). They took the exact same pile of scientific documents—think physics papers, math textbooks, and technical reports—and split them into two groups.
- Group A (The Old Way): They converted the pages into plain text, stripping out all the visuals.
- Group B (The New Way - VP): They fed the raw images of the pages directly to the model, letting it "see" the figures, tables, and layouts without ever turning them into words first.
The result? The robot trained on the raw images (Group B) consistently outsmarted the one trained on plain text (Group A). In fact, on tricky science and math benchmarks, the visual learner scored higher. For example, on a tough math test called AIME, the visual model improved its score significantly more than the text model. The paper suggests that the "lossy" process of converting a complex page with a diagram into a string of words actually deletes the very information needed to solve hard problems.
The Efficiency Hack: Doing More with Less
Here's the really cool part: the visual method wasn't just better; it was also way more efficient. The text-based model had to chew through about 80 billion text tokens to learn from these documents. The visual model, however, only needed about 20 billion visual tokens to get the same (or better) results. That's using only 25% of the data budget!
Think of it like this: If the text model is trying to describe a city by listing every single street name and building address, it takes forever. The visual model, by contrast, just looks at a map. It gets the whole picture instantly. The researchers found that the more "visual stuff" (like complex diagrams and equations) a page had, the bigger the advantage the visual model had. On pages packed with charts and formulas, the visual model's advantage grew huge, while on plain text pages, they were about the same.
No Cheat Codes, Just Pure Vision
You might wonder, "Did they cheat by showing the robot pictures and the answers?" No. They didn't use any special "image-to-text" labels or cheat sheets. The robot was just told to look at the next patch of the image and guess what it would look like, similar to how it guesses the next word in a sentence.
Surprisingly, this "just looking" training made the robot better at everything, not just looking. It got better at understanding text too, and even better at connecting pictures to words when tested later. The researchers measured this by checking how well the robot's "brain" aligned pictures with words, and found that the visual training made the two concepts fit together much more tightly, as if the robot finally understood that a picture of a cat and the word "cat" are actually the same thing.
What This Means (and What It Doesn't)
The authors are careful to say this isn't a magic wand that replaces reading text. Text is still great for learning facts that are already written down clearly. But for scientific documents, where the layout, the equation's shape, and the diagram's position carry crucial meaning, the visual approach suggests a new, scalable path forward.
They admit they haven't solved everything yet. For instance, they aren't sure if this works as well for random photos of nature or videos, since their tests were mostly on high-density scientific PDFs. But for the specific case of learning from complex documents, the paper suggests that letting the AI see the world as it is—visual, messy, and full of diagrams—might be the key to unlocking its true intelligence. It's a shift from "reading the description of the map" to "actually looking at the map."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.