← Latest papers
💻 computer science

Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force

This paper proposes MuSe, a multisensory continual learning framework that effectively adapts pretrained vision-only robot policies to new force-torque modalities using multi-stage fusion and experience replay, thereby enhancing performance on contact-rich tasks without forgetting original capabilities.

Original authors: Jaden Clark, Changhao Wang, Yihuai Gao, Seongheon Hong, Hojung Choi, Mark Cutkosky, Yifan Hou, Shuran Song

Published 2026-07-01
📖 5 min read🧠 Deep dive

Original authors: Jaden Clark, Changhao Wang, Yihuai Gao, Seongheon Hong, Hojung Choi, Mark Cutkosky, Yifan Hou, Shuran Song

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have trained a robot to be a master chef, but you only taught it using eyes. It learned to chop, stir, and plate food by watching thousands of videos. It's great at seeing what to do, but it doesn't know how hard to press or when something is slipping.

Now, imagine you want to teach this robot a new skill: feeling. You want it to use a force sensor (like a sensitive touch) to handle delicate tasks, like wiping a vase without breaking it or screwing a lightbulb in without cracking it.

Here's the problem: You don't have a massive library of videos that include both the robot's eyes and its new "touch" sensors. You only have a tiny, new dataset with these new sensors. If you just try to teach the robot with this small new data, it might get confused and forget everything it learned about seeing. This is called "catastrophic forgetting."

This paper introduces a solution called MuSe (Multisensory Continual Learning). Think of MuSe as a clever study guide that helps the robot learn a new sense without forgetting its old ones.

Here is how MuSe works, using simple analogies:

1. The "Two-Way Street" Connection (Multi-Stage Fusion)

Usually, when you add a new sense to a robot, you just tack it on at the end, like adding a side dish to a meal. MuSe is different. It creates a two-way street between the robot's eyes and its new touch sensors.

  • Early Fusion: It mixes the "touch" data with the "sight" data right at the beginning, like mixing flour and eggs before baking. This lets the robot understand how touch and sight relate to each other immediately.
  • Late Fusion: It also checks in later in the process, like a taste-tester adjusting the seasoning halfway through cooking.
  • The Result: The robot doesn't just "see" or "feel"; it learns a unified understanding where sight and touch help each other.

2. The "Future Crystal Ball" (Multisensory Future Prediction)

To make sure the robot really understands the new sense, MuSe doesn't just ask it to "do the task." It asks the robot to predict the future.

  • The robot is trained to guess: "What will I see in the next second? What will I feel in the next second? What action should I take?"
  • The Analogy: Imagine a driver learning to drive in the rain. Instead of just driving, they are asked to predict, "If I turn the wheel now, will the car skid?" and "Will the windshield wipers clear the water?"
  • By forcing the robot to predict both the visual scene and the force signals, it builds a deep, internal map of how the physical world works. This helps it generalize, meaning it can use its new "touch" skills even on tasks it hasn't seen before.

3. The "Flashcard Review" (Experience Replay)

This is the most critical part for preventing the robot from forgetting. When the robot learns the new "touch" skills, MuSe forces it to keep practicing its old "sight" skills at the same time.

  • The Analogy: Imagine a student learning a new language (French) while trying to keep their Spanish skills sharp. If they only study French for a month, they might forget Spanish. MuSe is like a teacher who says, "For every hour you study French, you must spend 30 minutes reviewing Spanish."
  • In the robot's case, it mixes the tiny new "touch" data with the huge old "sight" data. When the robot sees a task where it doesn't have touch data (because the old data didn't have it), MuSe simply covers up the "touch" part of the question and lets the robot answer based on sight alone. This keeps the old skills alive.

What Did They Find?

The researchers tested this on a real robot arm.

  • Forward Transfer (Learning New Things): The robot became much better at contact-heavy tasks (like wiping a vase or inserting a peg) using the new touch sensors. It didn't just learn the task; it learned how to feel its way through it.
  • Backward Transfer (Not Forgetting): Surprisingly, after learning to use the touch sensors, the robot actually got better at its original vision-only tasks. It seems that learning to "feel" helped it understand the physics of the world better, which made its "sight" skills sharper too.
  • Cross-Modal Generalization: Even on tasks where they never showed the robot the touch data during training, the robot could still predict what the forces would have been. It learned a general concept of "force" that it could apply anywhere.

The Bottom Line

MuSe shows that you don't need a massive, perfect dataset of every possible sensor to train a robot. You can take a robot that is already an expert at seeing, give it a small amount of new "touch" data, and use this special training method to upgrade it. The robot learns the new sense, keeps its old skills, and actually becomes smarter overall.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →