← Latest papers
💻 computer science

MoRI: Mixture of RL and IL Experts for Long-Horizon Manipulation Tasks

MoRI is a novel robotic manipulation framework that dynamically combines Imitation Learning and Reinforcement Learning experts to achieve high success rates, rapid convergence, and significantly reduced human intervention in complex long-horizon tasks.

Original authors: Yaohang Xu, Lianjie Ma, Gewei Zuo, Wentao Zhang, Han Ding, Lijun Zhu

Published 2026-04-15
📖 5 min read🧠 Deep dive

Original authors: Yaohang Xu, Lianjie Ma, Gewei Zuo, Wentao Zhang, Han Ding, Lijun Zhu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to perform a complex task, like folding a towel or putting a block into a drawer. You have two main ways to teach it:

  1. The "Copycat" (Imitation Learning): You show the robot a video of a human doing the task perfectly, and the robot tries to copy it.

    • Pros: It learns fast.
    • Cons: If the robot makes a tiny mistake early on, it gets confused and keeps making bigger mistakes. It's like a student who memorizes the answers but doesn't understand the math; if the test question changes slightly, they fail.
  2. The "Explorer" (Reinforcement Learning): You let the robot try, fail, and try again, giving it a "gold star" (reward) when it succeeds.

    • Pros: It learns to handle surprises and can figure out new ways to do things.
    • Cons: It takes forever. It might crash into things a thousand times before it learns the right move. It's like a toddler learning to walk by falling down repeatedly.

The Problem

Most real-world tasks need both. You need the Copycat to handle the easy, predictable parts (like moving your arm from point A to point B), and you need the Explorer to handle the tricky, messy parts (like figuring out exactly how to grip a slippery towel or slide a plug into a tight socket).

Existing methods tried to mix these two, but they often got confused about when to use which strategy, leading to slow learning or the robot needing a human to constantly step in and fix it.

The Solution: MoRI (Mixture of RL and IL Experts)

The authors of this paper created a system called MoRI. Think of MoRI not as a single robot brain, but as a team of two specialists working together under a smart manager.

The Team Members

  1. The Veteran (IL Expert): This is the "Copycat." It's great at smooth, routine movements. It knows exactly how to open a drawer or move a block because it has seen humans do it before.
  2. The Adventurer (RL Expert): This is the "Explorer." It's good at figuring out the tricky, high-precision stuff where things might go wrong. It's willing to take risks to find the perfect way to grip a slippery object.

The Manager (The Gating Network)

This is the magic part. MoRI has a smart manager that watches the two experts. It asks: "How sure are we about the next move?"

  • If the experts agree (low variance): The manager says, "This is easy. Let the Veteran do it." (e.g., "Just move the arm to the drawer.")
  • If the experts disagree or are unsure (high variance): The manager says, "This is tricky. Let the Adventurer take over and figure it out." (e.g., "How do I grip this towel without it slipping?")

This switching happens instantly, like a relay race where the baton is passed exactly when the runner needs to sprint.

How They Trained It (The Two-Stage Plan)

To make sure the team works well without breaking the robot, they used a two-step training process:

  1. Stage 1: The Classroom (Offline Pre-training):
    Before the robot ever touches a real object, they fed it 20 videos of humans doing the tasks.

    • The Veteran memorized these videos.
    • The Adventurer studied them to get a head start, so it doesn't have to start from zero.
    • Analogy: It's like studying a textbook before taking a driving test.
  2. Stage 2: The Driving Range (Online Fine-tuning):
    Now the robot goes to the real world.

    • The Manager directs the traffic.
    • If the robot gets stuck, a human gives a tiny nudge (intervention).
    • Crucially, the system uses a "safety net" (Regularization). It tells the Adventurer: "You can explore, but don't go too far off the path the Veteran showed you." This keeps the robot safe and prevents it from crashing while learning.

The Results: Why It's a Big Deal

The team tested this on four difficult real-world tasks (putting blocks in drawers, folding towels, inserting plugs, etc.). Here is what happened:

  • Speed: The robot learned in 2 to 5 hours. Other methods took much longer.
  • Success Rate: It succeeded 97.5% of the time.
  • Human Help: It needed human help 85% less than other methods. The robot figured out most of the problems on its own.
  • Smoothness: Because the Manager switches experts so smoothly, the robot's movements didn't jerk or jump around.

The Bottom Line

MoRI is like hiring a chess grandmaster (the Veteran) and a gambler (the Adventurer) and putting a referee (the Manager) between them. The grandmaster handles the standard moves, the gambler handles the wild cards, and the referee makes sure they switch roles at the perfect moment.

This allows robots to learn complex, long tasks quickly, safely, and with very little human babysitting, bringing us one step closer to robots that can actually help us around the house or in factories.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →