REPAIR-Bench: A Benchmark for Robot Error Perception And Interaction Recovery
REPAIR-Bench is a new benchmark derived from 214 human-robot interaction trials that addresses limitations in prior failure modeling by providing synchronized multimodal data and three novel evaluation tasks to advance the detection of inter-dependent failures, visual classification of failure types, and user-centered prediction of recovery strategies.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are working with a robot assistant in a busy hospital. You ask the robot to fetch a specific medicine, but something goes wrong. Maybe the robot gets confused, speaks the wrong words, or takes too long.
REPAIR-Bench is a new "training manual" and "exam" designed to teach robots how to realize when they've messed up and, more importantly, how to fix it in a way that makes the human feel better.
Here is a simple breakdown of what the researchers did, using everyday analogies:
1. The Problem: Robots Don't Know They're Failing
Think of a robot like a new employee who is very polite but doesn't realize they are making mistakes.
- Old Way: Previous studies treated every mistake as a single, isolated event (like a "pop quiz") and only asked, "Did it fail? Yes or No?" They also relied on rigid, pre-written rules for how the robot should apologize.
- The New Approach: The researchers realized that humans react to mistakes in complex ways. If a robot fails once, the human gets a little annoyed. If it fails three times in a row, the human gets frustrated. The robot needs to understand this history and read the human's subtle body language (like a frown or a confused look) before the human even says a word.
2. The Dataset: The "Crash Cart" Simulation
To build this, the team created a special dataset called RFM-HRI.
- The Scene: They set up a simulation where people had to use a "crash cart" (a rolling medical cart with emergency supplies) guided by a robot.
- The Glitches: The researchers intentionally broke the robot in four specific ways:
- Speech: The robot said the wrong thing.
- Timing: The robot was too slow.
- Comprehension: The robot didn't understand the request.
- Search: The robot couldn't find the item.
- The Data: They recorded 41 people doing this task 5 times each. They didn't just record audio; they recorded facial expressions (using a tool that tracks tiny muscle movements), head movements, and eye gaze. They also asked the participants afterward: "How did you feel?" and "How would you have liked the robot to fix this?"
3. The Three "Exams" (Tasks)
The paper introduces three specific challenges to test if a robot is smart enough to handle these situations:
Task 1: The "Spot the Glitch" Timer
- The Goal: Can the robot look at your face and say, "Oh no, I just messed up," at the exact second it happened?
- The Trick: The robot has to look at your face before you say anything. It's like a waiter noticing you look confused the moment they drop a tray, before you even yell "Hey!"
- The Result: They built a "Hierarchical" brain (a computer model that remembers past interactions). This model was much better at spotting the failure than models that only looked at the current moment. It could pinpoint the failure within about 3 seconds.
Task 2: The "Diagnosis" Class
- The Goal: Can the robot tell what kind of mistake it made just by looking at you?
- The Analogy: If a doctor sees a patient coughing, they need to know if it's allergies, a cold, or something else. Similarly, the robot needs to know: "Did I speak too fast (Timing), or did I say the wrong drawer number (Search)?"
- The Result: The robot got pretty good at telling if it was a "Success" or a "Speech" error, but it still struggled to distinguish between "Timing" and "Comprehension" errors. It's like a student who knows they failed the test but isn't sure if they failed because they didn't study or because the questions were tricky.
Task 3: The "Apology" Generator
- The Goal: Once the robot knows it failed, what should it say or do?
- The Analogy: Imagine you spill coffee on a friend. You could say "Sorry," "Let me get a towel," or "I'll buy you a new shirt." Different people prefer different fixes. The robot needs to guess what you specifically want.
- The Method: They used a type of AI called a "Small Language Model" (like a smart chatbot) and "fine-tuned" it (taught it specifically on this data) to suggest the best recovery strategies.
- The Result: The fine-tuned robot was much better at guessing the user's preferred fix than the untrained version. It learned that for some errors, an apology works, but for others, the user just wants a step-by-step guide to finish the task.
4. Why This Matters
The paper argues that for robots to be trusted in high-stakes places like hospitals, they can't just be "dumb machines" that crash and wait for a human to reboot them. They need to:
- Notice the failure through body language.
- Understand the type of failure.
- Adapt their recovery strategy to what the human actually wants.
The researchers found that by giving the robot a "memory" of past interactions and training it to read subtle facial cues, it becomes much more reliable. They also found that simply giving the robot a list of rules isn't enough; it needs to learn from human feedback to know how to apologize or fix things correctly.
In short: REPAIR-Bench is a new tool that helps robots learn to be better teammates by teaching them to read the room, admit their mistakes, and fix them in a way that keeps humans calm and trusting.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.