← Latest papers
💻 computer science

Towards Dexterous Embodied Manipulation via Deep Multi-Sensory Fusion and Sparse Expert Scaling

This paper introduces DeMUSE, a deep multimodal framework that integrates RGB, depth, and 6-axis force data via a Diffusion Transformer with adaptive normalization and sparse expert scaling to achieve state-of-the-art dexterous embodied manipulation in both simulation and real-world environments.

Original authors: Yirui Sun, Guangyu Zhuge, Keliang Liu, Jie Gu, Zhihao xia, Qionglin Ren, Chunxu tian, Zhongxue Ga

Published 2026-02-24
📖 5 min read🧠 Deep dive

Original authors: Yirui Sun, Guangyu Zhuge, Keliang Liu, Jie Gu, Zhihao xia, Qionglin Ren, Chunxu tian, Zhongxue Ga

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to perform delicate tasks, like pouring a glass of cola without spilling it, or sweeping crumbs into a dustpan without knocking the dustpan over.

For a long time, we've tried to teach robots mostly by showing them videos. We say, "Look at the picture, and do what you see." This works okay for simple things, but it's like trying to drive a car while wearing blindfolds and only looking at a rearview mirror. You miss the feeling of the steering wheel, the vibration of the road, and the pressure on the brakes.

This paper introduces a new robot brain called DeMUSE. Think of DeMUSE not just as a pair of eyes, but as a super-sensory brain that combines sight, touch, and "feeling" into one unified thought process.

Here is how it works, broken down with simple analogies:

1. The Problem: The "Blind" Robot

Current robots are like vision-only chefs. If you ask them to crack an egg, they look at the egg. But if they squeeze too hard, they crush it. If they squeeze too soft, nothing happens. They don't "feel" the shell cracking. They lack force feedback (how hard they are pushing) and depth perception (exactly how far away things are).

2. The Solution: The "Super-Sensory" Brain

DeMUSE solves this by feeding the robot three types of information at once, all mixed together like ingredients in a smoothie:

  • RGB (Eyes): What the robot sees (colors, shapes).
  • Depth (3D Vision): How far away objects are (like a 3D map).
  • 6-Axis Force (Touch): How hard the robot is pushing or pulling (like the feeling in your fingertips).

Instead of treating these as separate streams (one for eyes, one for hands), DeMUSE smashes them all together into a single "stream of consciousness." This allows the robot to understand that seeing a soft sponge and feeling it squish are part of the same event.

3. The Secret Sauce: "Adaptive Normalization" (The Volume Knob)

Here is a tricky part: The data from the camera is huge (millions of pixels), but the data from the force sensor is tiny (just a few numbers). If you mix them directly, the camera data drowns out the force data, like trying to hear a whisper while a jet engine is roaring.

DeMUSE uses a clever trick called AdaMN (Adaptive Modality-specific Normalization).

  • The Analogy: Imagine a sound mixer at a concert. The camera is the loud drums, and the force sensor is a quiet violin. A normal mixer would just turn everything up to the same volume, making the violin sound distorted.
  • DeMUSE's Mixer: It has a smart engineer who listens to the music and automatically adjusts the volume knob for each instrument separately. It turns down the drums just enough so the violin can be heard clearly, without losing the rhythm. This ensures the robot pays attention to the "touch" just as much as the "sight."

4. The Engine: "Sparse Experts" (The Specialized Team)

To make the robot smart enough to handle complex tasks, you usually need a massive computer brain. But massive brains are slow, and robots need to react instantly (in milliseconds).

DeMUSE uses a Mixture-of-Experts (MoE) architecture.

  • The Analogy: Imagine a giant hospital. Instead of having one doctor who tries to know everything about every disease (which is slow and tiring), you have a team of specialists.
    • When the robot sees a slippery object, it calls the "Friction Expert."
    • When it needs to stack blocks, it calls the "Gravity Expert."
    • When it needs to push a button, it calls the "Force Expert."
  • The Magic: Only the specific expert needed for the job wakes up and does the work. The rest of the team sleeps. This makes the robot incredibly smart (it has a huge "team" of knowledge) but incredibly fast (it only uses a small amount of energy at any given moment).

5. The Training: "Dreaming" the Future

DeMUSE is trained using a method called Diffusion.

  • The Analogy: Imagine the robot is learning to paint. At first, it sees a canvas covered in static noise (random pixels). It has to guess what the final picture should look like.
  • DeMUSE doesn't just guess the picture; it guesses the future. It asks, "If I move my hand this way, what will the world look like in 0.5 seconds?" It practices "dreaming" the future outcome and the physical movement simultaneously. This ensures that what the robot imagines happening matches what actually happens when it moves.

The Results: From Simulation to Reality

The researchers tested DeMUSE on tough tasks:

  • Filling a cup with exactly 1/3 of a cola (requires precise stopping).
  • Sweeping crumbs into a dustpan (requires coordinating a broom and a pan).
  • Organizing tools in a drawer (requires feeling the drawer slide).

The Outcome:

  • In the computer simulation, it succeeded 83% of the time.
  • In the real world (with real robots and real gravity), it succeeded 72.5% of the time.

This is a huge jump compared to previous robots, which often failed at these tasks because they couldn't "feel" their way through the problem.

Summary

DeMUSE is a robot brain that finally learns to see, feel, and think all at once. By treating touch and sight as equal partners, using smart volume knobs to balance them, and employing a team of specialized "experts" to keep things fast, it can perform delicate, human-like tasks that were previously impossible for machines. It's the difference between a robot that just looks at a task and a robot that truly understands it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →