How Users Understand Robot Foundation Model Performance through Task Success Rates and Beyond
This paper investigates how non-expert users interpret Robot Foundation Model performance metrics, finding that while they effectively utilize Task Success Rates, they also highly value failure case details and desire access to both historical evaluation data and the robot's own self-estimates for novel tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you just bought a brand-new, super-smart kitchen robot. You want it to put a vase on a high shelf. But here's the catch: you've never seen this specific robot do this specific task before. How do you know if it's going to succeed, or if it's going to knock the vase off the counter and shatter it?
This paper is like a "User Manual for Trust." The researchers wanted to find out how regular people (non-experts) understand the "report cards" that robot scientists give to these machines.
The "Report Card" Problem
Right now, when scientists test these robots, they mostly give them a single grade: Task Success Rate (TSR).
- The Analogy: Think of TSR like a batting average in baseball. If a robot has an 80% success rate, it means it got the job done 8 times out of 10.
- The Finding: The study found that regular people understand this number just fine. If the number is high, people feel confident letting the robot work alone. If the number is low, they feel nervous and want to watch closely.
The Missing Pieces of the Puzzle
However, the researchers discovered that a batting average isn't enough. Imagine if a baseball coach only told you, "He hits 80% of the time," but never told you how he fails.
- The Analogy: If you knew the player always trips over his own feet when running to first base, you'd know to stand back. But if you only knew the average, you might get hit by a ball.
- The Finding: The study showed that people really want to know about the failures. They want to hear a plain-language story like, "The robot usually succeeds, but sometimes it grabs the wrong bottle." This helps people prepare for the worst-case scenario.
The "Real Data" vs. "The Robot's Guess"
The researchers tested two different ways of giving information:
- Real Data (The Past): "We tried this exact task 5 times, and it worked 4 times."
- The Robot's Guess (The Estimate): "I think I can do this task 80% of the time."
The Surprise: People loved both, but they had a slight preference for Real Data.
- The Analogy: It's like buying a used car. You trust a mechanic's report saying, "This car has driven 50,000 miles without breaking down" (Real Data) more than the car's computer guessing, "I feel like I'll run fine today" (Estimate).
- However, people still found the robot's own estimate very useful, especially for tasks the robot had never tried before.
The "In-Person" Factor
The researchers did two studies: one where people watched videos on a computer, and one where they stood in a lobby watching a real robot try to put soup cans on a shelf.
- The Finding: Seeing the robot in real life didn't change what information people wanted. They still wanted the success rates and failure stories.
- The New Twist: When standing next to the robot, people got a little more worried about physical safety. They asked, "If it drops the can, will it break my floor?" or "Will it hit me?" They wanted to know about the robot's strength and speed, things they didn't think about as much when just watching a video.
The Bottom Line
The paper concludes that to make these robots safe and useful in our homes, we can't just show them a number (like "80%"). We need to give users a full "dossier" that includes:
- The Score: How often it succeeds.
- The Horror Stories: Clear descriptions of how it might fail.
- The Evidence: Real videos of it doing similar tasks.
- The Forecast: The robot's own honest guess about how it will do on a new task.
By giving people this full picture, they can make smarter decisions about when to let the robot work alone and when to stand by with a safety net.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.