From Action Labels to Sets: Rethinking Action Supervision for Imitation Learning from Corrective Feedback
This paper introduces CLIC, a human-in-the-loop imitation learning framework that replaces brittle pointwise action supervision with set-valued targets derived from corrective feedback, enabling robust policy learning that accommodates noisy, relative, and multi-modal human guidance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to perform a task, like pouring water into a cup or catching a ball. The traditional way to do this is called Behavior Cloning. Think of it like a strict teacher who points to a single, exact spot on a map and says, "You must go exactly here."
The problem with this approach is that humans aren't perfect. Sometimes we get tired, our hands shake, or we just give a slightly wrong direction. If the robot treats every single human instruction as a perfect, unchangeable law, it starts to memorize our mistakes. It becomes "brittle"—if the human makes a tiny error, the robot gets confused and fails. It's like a student who memorizes the exact coordinates of a wrong answer on a test and fails because they can't handle any variation.
On the other hand, some methods try to teach the robot by saying, "This action is better than that one." This is like a teacher saying, "Go left is better than going right," without saying exactly how far left to go. While this is more flexible, it's often too vague. The robot might end up wandering aimlessly because it doesn't know the boundaries of "good" behavior.
The New Idea: "The Safe Zone"
The authors of this paper propose a middle ground called CLIC (Contrastive policy Learning from Interactive Corrections). Instead of giving the robot a single point or a vague comparison, they teach it to look for a "Safe Zone" or a Target Set.
Here is the analogy:
- Old Way (Pointwise): The teacher says, "Stand exactly on this specific tile." If the teacher points to a slightly wrong tile, the robot stands there and fails.
- Old Way (Pairwise): The teacher says, "Standing on the left side is better than the right side." The robot knows it shouldn't be on the right, but it doesn't know if it should be on the far left, the middle, or just slightly left. It's too loose.
- CLIC (Set-Valued): The teacher says, "Don't stand on the red tile (where the robot messed up), but you can stand anywhere in this blue circle around the green tile."
In this new method, when a human corrects the robot, they aren't just giving a new coordinate. They are defining a region of acceptable actions.
- If the human pushes the robot's arm slightly to the left to correct a mistake, the robot learns: "Okay, the correct action is somewhere in this general direction, not necessarily exactly where you pushed me."
- If the human gives a noisy or shaky correction, the robot understands that the "Safe Zone" is a bit fuzzy, but it still knows what not to do (the wrong side) and what is generally acceptable (the zone).
How It Works in Practice
The paper tested this idea in two main ways:
Simulations: They ran thousands of computer simulations with tasks like pushing a T-shaped block, picking up a can, or lifting an object with two robot arms.
- The Result: When the human data was perfect, CLIC worked just as well as the old methods. But when the human data was noisy, shaky, or only partial (e.g., the human only corrected one arm of a two-armed robot), CLIC kept working while the old methods crashed or failed completely.
- Why? Because CLIC didn't try to memorize the shaky hand movements. It learned the shape of the correct behavior.
Real Robots: They tested this on actual robots doing tasks like:
- Inserting a T-shape into a U-shape: A complex puzzle requiring precise movement.
- Catching a ball: A fast, dynamic task where humans can't easily give perfect demonstrations, so they just give quick directional nudges.
- Pouring water: A task requiring delicate control of a bottle.
- The Result: The robot learned faster and more reliably. In the ball-catching task, the robot went from rarely catching the ball to catching it consistently within two tries, even though the human was only giving rough directional hints.
The Key Takeaway
The paper claims that by treating human corrections as regions of possibility rather than exact targets, robots become much more robust. They can handle human mistakes, shaky hands, and incomplete instructions without getting confused.
It's the difference between a student who memorizes a single wrong answer and a student who understands the concept of the right answer. CLIC teaches the robot the concept (the "Safe Zone") rather than just the coordinates, making it a much better learner when humans are involved.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.