← Latest papers
💻 computer science

Inverse Manipulation through Symbolic Planning and Residual Operator Learning

This paper proposes a hybrid framework for robotic inverse manipulation that combines automatically extracted symbolic operators with residual Reinforcement Learning to refine coarse symbolic plans into physically grounded skills, demonstrated effectively on the ManiSkill3 PushCube task.

Original authors: Yigit Yildirim, Giuseppe Rauso, Riccardo Caccavale, Alberto Finzi

Published 2026-06-05
📖 4 min read☕ Coffee break read

Original authors: Yigit Yildirim, Giuseppe Rauso, Riccardo Caccavale, Alberto Finzi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot how to push a cube across a table. That's the "forward" task. Now, imagine you need to teach that same robot how to undo that exact action and get the cube back to where it started.

This sounds simple, but it's actually a tricky puzzle for a robot. If you just tell the robot to "play the movie backward," it often fails. Why? Because if the robot didn't grab the cube (it just pushed it), it can't simply pull its arm back to undo the push. The cube is now sitting somewhere else, and the robot needs a new plan to get it back.

This paper presents a clever "hybrid" system that solves this problem by combining logical planning (like a chess player thinking ahead) with trial-and-error learning (like a baby learning to walk).

Here is how their system works, broken down into simple steps:

1. The "Magic Translator" (Symbolic Extraction)

First, the robot watches a human (or a script) perform the forward task, like pushing a cube. Instead of just memorizing the arm movements, the system acts like a translator. It watches the scene and writes down a simple "recipe" or rule in plain logic.

  • Before: "The cube is at the start."
  • After: "The cube is at the finish."
  • The Rule: "If the cube is at the start, I can push it to the finish."

2. The "Reverse Recipe" (Inverse Planning)

Once the robot has the rule, it automatically writes a "reverse recipe." It looks at the rule and asks, "What do I need to do to get back to the start?"

  • It realizes it needs to: "Put the cube back at the start," "Make sure the gripper is open," and "Make sure the robot arm is near the cube."
  • The robot then uses a standard planner (like a GPS) to figure out a sequence of basic moves to try and fix the situation. For example, it might try to "Pick up the cube" and "Place it at the start."

3. The "Good Enough" Handoff

Here is where the magic happens. The robot's basic planner is good, but not perfect. It might successfully pick up the cube and drop it roughly near the start, but it might be off by a few inches.

  • In many other systems, this would be a failure.
  • In this system, the robot stops and says, "Okay, I got the cube close enough. Now, I need to fine-tune the position."

4. The "Fine-Tuner" (Residual Learning)

This is the second half of the team: a Reinforcement Learning (RL) agent. Think of this as a highly skilled micro-adjuster.

  • The system tells the RL agent: "You must move the cube the remaining distance to the exact starting spot, but do not mess up the things we already fixed (like keeping the gripper open)."
  • The RL agent tries thousands of tiny movements. If it moves the cube closer, it gets a "treat" (a reward). If it knocks the cube away from the spot we already fixed, it gets a "scolding" (a penalty).
  • Eventually, the RL agent learns the perfect, tiny nudges needed to slide the cube exactly into place.

The Result: A Perfect Undo

The authors tested this on a task called "PushCube."

  • The Problem: The robot pushed a cube.
  • The Old Way (Just reversing): Couldn't do it because the robot didn't grab the cube.
  • The "Logic Only" Way: Got the cube close, but missed the target by about 17 millimeters (about the width of a pencil eraser).
  • The New Hybrid Way: The logic part got the cube close, and the learning part nudged it the rest of the way. The result? The cube ended up within 1.4 millimeters of the start, with a 90% success rate.

The Big Picture

The paper claims that by splitting the job into two parts—Symbolic Planning for the big picture and Residual Learning for the fine details—you can teach robots to undo complex tasks that they couldn't reverse before. It turns a "rough guess" into a "physically perfect undo."

In short: The robot uses its brain to figure out the general strategy to undo a task, and then uses its "muscle memory" (learned through trial and error) to make the final, precise adjustments to get it exactly right.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →