Position: Good Embodied Reward Models Need Bad Behavior Data
This position paper argues that the embodied AI community must prioritize collecting and utilizing "bad" robot data—such as failed, suboptimal, or hazardous behaviors—to train reward models that accurately align with human preferences and avoid over-rewarding unsafe or superficial task completions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to fold laundry or pour a glass of water. To teach it well, you need a "teacher" (a reward model) that can look at what the robot does and say, "Good job!" or "No, that was a mess."
This paper argues that our current robot teachers are failing because they have only ever seen perfect examples. They have never been shown what a disaster looks like.
Here is the breakdown of the paper's argument using simple analogies:
1. The Problem: The "Perfect Student" Bias
Currently, when we train these robot teachers, we only show them videos of robots doing tasks perfectly. It's like trying to teach a student how to drive a car by only showing them videos of professional drivers in perfect weather, never showing them a flat tire, a skid, or a near-miss.
Because the teachers have never seen "bad" behavior, they are overly optimistic.
- The Result: If a robot knocks over a bowl while trying to pick up a cup, the teacher might still say, "Great job! You picked up the cup!" because it sees the cup moving. It misses the fact that the robot broke something else.
- The Reality: The paper tested three top-of-the-line robot teachers. They all gave high scores to robots that were actually failing, being unsafe, or taking "cheats" to finish a task. As tasks got harder (like using a tool), the teachers got worse, essentially guessing randomly.
2. The Solution: We Need "Bad" Data
The authors say we need to stop throwing away the "failures." We need to collect and share videos of robots crashing, dropping things, or moving dangerously.
Think of it like a medical school. If you only study healthy patients, you won't know how to diagnose a disease. You need to study sick patients, too.
- The Paper's Claim: Even a small amount of "bad" data (videos of robots failing) helps the teacher learn what not to reward.
- The Experiment: The researchers tested this by showing a teacher a video of a robot failing (e.g., spilling nuts) alongside a text description of the error.
- Text only: Didn't help much. The teacher couldn't connect the words to the physical mess.
- Text + Video: Helped a little on simple tasks (like picking up an object).
- Text + Video + "Scorecard": This worked best. They gave the teacher a video of the failure plus a detailed scorecard showing exactly when and why the robot failed during the video. This allowed the teacher to learn to penalize the robot the moment it made a mistake, even if the rest of the task looked okay.
3. Why We Don't Have This Data Yet
You might ask, "Why don't we just record all the failures?"
- The Barrier: In the digital world (like text chatbots), generating "bad" data is easy and cheap. You can just type nonsense.
- The Robot World: In the real world, robots are physical. Making them fail often means breaking things, wasting time, or creating safety hazards. Also, most robotics labs are so focused on showing off success that they delete the "ugly" failure videos before anyone sees them.
4. The Call to Action
The authors are asking the robotics community to do four specific things:
- Release the "Bad" Data: Stop hiding the failures. Labs and companies should share datasets of robots crashing, dropping things, or moving unsafely, just as they share success stories.
- Fake the Failures (Synthetically): Since real crashes are expensive, we need better computer simulations that can generate realistic "what-if" failure scenarios (e.g., "What if the robot slips here?").
- Decentralize Testing: Instead of one lab testing robots, we need many different places testing them in real-world conditions to catch more types of failures.
- Better Tests for Teachers: We need new benchmarks to test the "teachers" themselves, ensuring they are actually learning from human feedback and not just guessing.
Summary
The paper argues that to build reliable robots, we must stop treating "failure" as a mistake to be hidden. Instead, we must treat failure data as a vital resource. Without seeing the "bad" behavior, our robot teachers will keep giving high scores to dangerous or sloppy robots, thinking they are doing a great job.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.