← Latest papers
🤖 AI

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection

The paper introduces RoCo-ACE, a rollout-conditioned online distillation framework that enhances knowledge injection in pretrained MLLMs by dynamically reallocating distillation weights to reference-supported tokens and applying sparse anchored corrections, thereby achieving superior injected-knowledge accuracy while effectively preserving the model's original behavioral retention.

Original authors: Yan Hong, Wei Li, Kedong Xiu, Jun Lan, Shuheng Zhou, Zhongcai Lyu, Huijia Zhu, Weiqiang Wang, Jianfu Zhang

Published 2026-07-29
📖 3 min read☕ Coffee break read

Original authors: Yan Hong, Wei Li, Kedong Xiu, Jun Lan, Shuheng Zhou, Zhongcai Lyu, Huijia Zhu, Weiqiang Wang, Jianfu Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot friend who has read almost every book in the library and seen millions of pictures. This robot is great at chatting, solving puzzles, and describing what it sees. But here's the catch: the world changes every day. New movies come out, new scientific discoveries happen, and famous people do new things. If you want your robot friend to know about these fresh facts, you have to teach it. But teaching a giant robot is tricky. If you just force it to memorize a new fact, it might forget how to do its old tricks, like drawing a cat or answering a safety question. It's like trying to teach a chess grandmaster a new opening move, but in the process, they forget how to play the game entirely. This paper tackles that exact problem: how to slip new, specific facts into a smart robot's brain without messing up its old, reliable skills.

The researchers behind this study, working with models from Ant Group and universities, realized that the usual ways of teaching these robots were a bit clumsy. One method is like forcing the robot to copy a textbook word-for-word; it learns the fact, but it also starts sounding like a boring robot and forgets its personality. Another method tries to be very careful and only tweak tiny parts of the brain, but it often misses the point, leaving the new facts half-learned. The team, led by Yan Hong and Jun Lan, came up with a clever new strategy called RoCo-ACE. Think of it as a "smart highlighter" system. Instead of making the robot memorize the whole textbook, they let the robot try to answer the question first. Then, they compare the robot's answer with the "correct" answer from a teacher.

Here is the magic trick: The system looks at the robot's answer and asks, "Did the teacher's help make this specific word more likely?" If the robot guessed a correct fact (like a specific date or name) and the teacher's presence made that guess even stronger, the system gives that word a high-five and tells the robot to remember it well. But if the robot is just using generic words (like "a big building" instead of "the Eiffel Tower"), the system ignores them. This is the RoCo part. Then, there's the ACE part. Sometimes, the robot completely misses the new fact. The ACE system acts like a gentle nudge, saying, "Hey, you forgot the name of the museum! Let's fix just that one missing piece." By combining these two, the robot learns the new facts accurately without losing its ability to chat naturally or stay safe.

The team tested this on three different types of knowledge updates, from news about celebrities to science facts. They found that RoCo-ACE was the best at teaching the new facts compared to other methods. In one test, it boosted the robot's knowledge accuracy to 27.6 (compared to 20.1 for the standard "copy the textbook" method). Crucially, while other methods caused the robot to forget its old skills (dropping its general performance score to as low as 31.8), RoCo-ACE kept the robot's general skills almost exactly where they started, at around 56.5. The paper suggests that this approach offers a much better balance: you get the new information without the robot losing its mind. It's a step toward robots that can keep learning new things every day without forgetting who they are.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →