← Latest papers
💻 computer science

Sumo: Dynamic and Generalizable Whole-Body Loco-Manipulation

This paper introduces Sumo, a sim-to-real approach that combines a pre-trained whole-body control policy with test-time sample-based planning to enable legged robots to dynamically and generally manipulate large, heavy objects across diverse tasks without additional training.

Original authors: John Z. Zhang, Maks Sorokin, Jan Brüdigam, Brandon Hung, Stephen Phillips, Dmitry Yershov, Farzad Niroui, Tong Zhao, Leonor Fermoselle, Xinghao Zhu, Chao Cao, Duy Ta, Tao Pang, Jiuguang Wang, Preston
Published 2026-04-10
📖 5 min read🧠 Deep dive

Original authors: John Z. Zhang, Maks Sorokin, Jan Brüdigam, Brandon Hung, Stephen Phillips, Dmitry Yershov, Farzad Niroui, Tong Zhao, Leonor Fermoselle, Xinghao Zhu, Chao Cao, Duy Ta, Tao Pang, Jiuguang Wang, Preston Culbertson, Zachary Manchester, Simon Le Cléac'h

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a robot dog named Spot. Now, imagine you want Spot to do something incredibly difficult: stand up a heavy tire that weighs more than Spot itself, or drag a giant crowd-control barrier that's taller than the robot.

Usually, teaching a robot to do this is like trying to teach a toddler to perform brain surgery. You have to tell the robot exactly how to move every single muscle (joint) at every single millisecond. If you get the math wrong, the robot falls over. If the object is slightly different (like a tire instead of a box), the robot gets confused and fails.

Enter "Sumo."

The paper introduces a new system called Sumo (Dynamic and Generalizable Whole-Body Loco-Manipulation). Think of Sumo not as a robot, but as a brilliant coaching strategy that lets the robot use its whole body to solve these tough puzzles.

Here is how it works, broken down into simple concepts:

1. The "Muscle Memory" vs. The "Coach"

To understand Sumo, you need to understand the two parts working together:

  • The Low-Level Policy (The Muscle Memory): Imagine the robot has already spent years training in a video game simulator. It has learned "muscle memory" for walking, balancing, and moving its limbs. It knows how to stand up without falling over. It's like a professional athlete who can run without thinking about how to move their legs.

    • In the paper: This is the pre-trained RL (Reinforcement Learning) policy. It handles the boring, hard math of keeping the robot upright.
  • The High-Level Planner (The Coach): This is the new part. The "Coach" doesn't tell the robot how to move its knees or elbows. Instead, the Coach gives high-level instructions like, "Push that tire forward," or "Grab that chair and lift it."

    • In the paper: This is the Sample-Based MPC (Model Predictive Control). It runs a mental simulation thousands of times a second to figure out the best goal for the robot to aim for.

The Magic: The Coach tells the Muscle Memory what to do, and the Muscle Memory figures out how to do it. This is much easier than telling the robot exactly how to move every joint.

2. The "Video Game" Analogy

Think of the robot like a character in a video game.

  • Old Way (End-to-End RL): You try to teach the character to play the level by pressing buttons randomly until they win. If the level changes (a new enemy appears), you have to start training from scratch. It takes forever.
  • Old Way (End-to-End MPC): You try to calculate the exact physics of every single pixel and button press in real-time. The computer gets overwhelmed and crashes because there are too many variables.
  • The Sumo Way: You give the character a pre-trained "movement pack" (they already know how to run and jump). Then, you use a smart AI coach to look at the map, say, "Okay, we need to jump over that wall and push that crate," and the character's muscle memory handles the rest. If the wall moves, the coach just updates the plan; the character doesn't need to relearn how to run.

3. Why is this a Big Deal?

The researchers tested this on a real robot dog (Boston Dynamics' Spot) and a simulated humanoid robot (Unitree G1). Here is what they achieved:

  • Lifting the Impossible: The robot lifted a 15kg tire (heavier than the robot's arm can usually lift) by using its legs, torso, and arm all at once. It was like a human doing a deadlift with their back, legs, and arms combined.
  • Generalization (The "One Size Fits All" Trick): This is the coolest part. They trained the system on a few tasks, but then they threw it into a room with completely new objects (a traffic cone, a heavy chair, a tire rack).
    • The Result: The robot didn't need to be retrained. The "Coach" just looked at the new object, realized, "Oh, that's a tire, I need to push it differently," and adjusted the plan instantly.
    • Analogy: Imagine you learn to drive a car. With Sumo, if you get into a truck, a boat, or a bicycle, you don't need to go back to driving school. You just tell your brain, "I'm driving a boat now," and your existing skills (balancing, steering) adapt automatically.

4. The "Sim-to-Real" Bridge

Robots are notoriously bad at moving from the computer simulation to the real world (the "Sim-to-Real Gap"). Real tires are bumpy; real floors are slippery.

  • Sumo bridges this gap because the "Coach" is constantly checking the real world. If the tire slips, the Coach sees it immediately and changes the plan for the next split second. It's like a driver adjusting their steering wheel the moment they feel the car skid, rather than trying to memorize the exact path of the road beforehand.

Summary

Sumo is a system that gives robots a "brain" (the planner) and a "body" (the pre-trained muscle memory).

  • Before: Robots were like clumsy toddlers who needed to be taught every single movement for every single object.
  • Now: Robots are like skilled athletes with a smart coach. They can look at a heavy, weirdly shaped object, figure out a plan on the fly, and use their whole body to move it, even if they've never seen that specific object before.

The paper proves that by combining learning (to get the muscle memory) with planning (to get the strategy), robots can finally start doing the kind of heavy, dynamic lifting and moving that humans do naturally.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →