← Latest papers
🤖 AI

Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression

This paper introduces CAPS, a cross-modal agentic policy self-distillation framework that bridges the performance gap between text-based and vision-compressed agent histories by using a model's stronger text-policy to supervise its visual-history counterpart, thereby significantly improving decision-making accuracy while drastically reducing memory context costs.

Original authors: Cheng Fan, Junyi Zhou, Tingzhang Luo, RongJian Xu, Qiyanhui Lu, Mingjian Zhu, Hanting Chen, Jianyuan Guo

Published 2026-08-11
📖 4 min read☕ Coffee break read

Original authors: Cheng Fan, Junyi Zhou, Tingzhang Luo, RongJian Xu, Qiyanhui Lu, Mingjian Zhu, Hanting Chen, Jianyuan Guo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a super-smart robot detective trying to solve a mystery. To do its job, it has to talk to a computer, ask questions, read the answers, and then decide what to do next. But here's the catch: every time the robot takes a step, it has to remember everything that happened before. If the mystery is long, the robot's "memory" gets huge, like a backpack that keeps getting stuffed with more and more heavy rocks. Eventually, the backpack gets so heavy that the robot moves slower, costs more money to run, and might even trip over its own thoughts.

Scientists have tried to solve this by taking all those heavy text memories and squishing them into a single, dense picture. It's like turning a 100-page novel into a single, information-packed comic book page. The idea is that the robot can look at the picture and remember the story without carrying the heavy text. But there's a problem: robots are usually trained to read words, not to "think" through pictures. When they switch from reading a book to looking at a comic, they often get confused. They might be able to read the words in the picture perfectly, but they forget how to use those words to solve the puzzle. They become great at reading, but terrible at reasoning.

This is exactly the puzzle a team of researchers from City University of Hong Kong and Huawei Technologies set out to solve. They discovered that simply turning text into images isn't enough; the robot needs to learn how to "think" in pictures, not just read them. They found that the gap between a robot that reads text and one that reads images isn't because the robot can't see the words (it can read them just fine). Instead, the robot loses its "game plan." It starts making bad choices, asking the wrong questions, or giving up too soon, even though it has all the information it needs right in front of its eyes.

To fix this, the team invented a clever training method called CAPS (Cross-modal Agentic Policy Self-distillation). Think of it like a master chef (the text-reading robot) teaching an apprentice (the image-reading robot) how to cook. The master chef doesn't just show the apprentice the ingredients; the chef watches the apprentice cook and says, "Hey, when you see this picture of the ingredients, you should chop the onions this way, not that way." The apprentice learns to mimic the master's decision-making process, even though the master is looking at a recipe book and the apprentice is looking at a photo of the ingredients.

The researchers tested this on two very different types of games: a search engine quiz where the robot has to find answers online, and a virtual house where the robot has to move objects around. They found that their new method, CAPS, made the image-reading robot much smarter. It didn't just get better at reading the pictures; it got better at reasoning with them. The robot started making the right choices, asking the right questions, and stopping at the right time, just like the text-reading robot.

The results were impressive. On the search quiz, the new method improved the robot's success rate by about 5% compared to previous methods. On the virtual house game, the improvement was even bigger, jumping up by over 15%. But the best part? The robot was still carrying that super-light "comic book" backpack. It solved the problems faster and used way less computer memory—cutting the memory cost by up to 83% in some cases.

The team showed that you don't have to choose between a heavy, slow memory and a fast, light one. By teaching the robot to think like a pro, even when looking at pictures, you can get the best of both worlds: a robot that is both incredibly smart and super efficient. They proved that the problem wasn't that the robot couldn't "read" the compressed history; the problem was that it needed to learn how to "reason" with it. And with CAPS, they finally taught it how to do just that.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →