CORE: Common Outcome Regularities from Action-Free Visual Demonstrations for Robot Manipulation
The paper introduces CORE, a robot imitation learning framework that leverages action-free human videos by extracting and utilizing common terminal outcome regularities as visual goal prototypes to overcome embodiment gaps and significantly improve manipulation success rates across simulated and real-world tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to stack bowls. You have two ways to give it instructions:
- The Text Manual: You tell the robot, "Please stack the bowls."
- The Video: You show the robot a video of a human stacking bowls, but you don't tell the robot how to move its arms to do it. You just let it watch the result.
For a long time, robots have struggled with the second option. They are great at following text instructions, but text is often too vague. "Stack the bowls" doesn't tell the robot exactly how high the stack should be, how the bowls should touch, or what the final shape needs to look like. It's like telling a chef to "make a cake" without showing them what a finished cake looks like.
On the other hand, robots have also struggled to learn from human videos because humans and robots look very different. A human uses hands; a robot uses a gripper. A human moves their whole body; a robot moves a specific arm. Trying to copy a human's exact hand movements usually confuses the robot.
The Big Idea: "The Finish Line, Not the Race"
The authors of this paper, CORE, realized something clever. They noticed that while different people might stack bowls in totally different ways (using different speeds, different angles, or different hand grips), they all end up with the exact same result: a neat, stable stack of bowls.
They call these shared results "Common Outcome Regularities."
Instead of trying to teach the robot how to move (the race), they decided to teach the robot what the finish line looks like (the result).
How CORE Works (The Three Steps)
Think of CORE as a smart teacher who watches a bunch of videos and then gives the robot a "Goal Picture."
Step 1: The Detective (Learning the Result)
The system watches many videos of people successfully stacking bowls. It ignores the messy middle parts of the video (the waving arms, the pauses) and focuses only on the very last second of the successful attempts. It learns to recognize the "fingerprint" of a successful stack. It's like a detective who only cares about the solved crime scene, not the journey to get there.Step 2: The Artist (Creating the Goal)
The system takes all those "successful ending" snapshots and blends them together to create one perfect, ideal "Goal Picture." This isn't just one random photo; it's a summary of what a perfect stack looks like, filtering out any weird or accidental ways people might have finished the task.Step 3: The Coach (Guiding the Robot)
Now, the robot tries to do the task. As it moves, the system constantly compares what the robot is seeing right now with that "Goal Picture." It tells the robot, "You are getting closer to the Goal Picture," or "You are moving away from it." This acts like a GPS that constantly corrects the robot's path until it matches the perfect ending.
Why This is Better Than Text
The paper tested this on many different tasks, like opening drawers, stacking blocks, and moving objects.
- Text Instructions are like a vague map: "Go to the park." The robot might get lost because it doesn't know exactly where the park bench is or how to climb the fence.
- CORE's Visual Goals are like a photo of the destination: "Stop exactly when you look like this."
In their tests, robots using CORE were much more successful than those using text instructions.
- In a simulation called Meta-World, robots improved their success rate by about 4%.
- In RoboTwin 2.0, they improved by 11%.
- In the Real World (using a physical robot arm), the improvement was huge: 17%.
The Real-World Proof
The researchers even tried this on a real robot arm in a lab. They asked the robot to open a drawer and stack three bowls.
- Robots using only text instructions often stopped just short of the target or fumbled the stack.
- Robots using CORE's "Goal Picture" knew exactly when they had reached the perfect position and completed the task successfully.
The Takeaway
The paper argues that we don't need to teach robots to copy human movements exactly. Instead, we should teach them to recognize what a "good job" looks like. By focusing on the outcome rather than the action, robots can learn from the vast amount of human videos available on the internet, even if they don't look like humans. It's a shift from asking "How do I move?" to asking "What does success look like?"
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.