← Latest papers
💻 computer science

ERR@HRI 3.0 Challenge: Multimodal Detection of Errors and Anticipation in Human-Robot Interactions

This paper presents the ERR@HRI 3.0 challenge, which introduced two naturalistic, crowdsourced video datasets to advance end-to-end multimodal machine learning for detecting bystander reactions to robot errors and anticipating potential failures, demonstrating that submitted models outperformed baseline systems.

Original authors: Maria Teresa Parreira, Micol Spitale, Maia Stiber, Shiye Cao, Amama Mahmood, Chien-Ming Huang, Hatice Gunes, Wendy Ju

Published 2026-07-14
📖 5 min read🧠 Deep dive

Original authors: Maria Teresa Parreira, Micol Spitale, Maia Stiber, Shiye Cao, Amama Mahmood, Chien-Ming Huang, Hatice Gunes, Wendy Ju

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're watching a robot try to make a sandwich. Sometimes, the robot drops the bread. Sometimes, it tries to put the toaster in the fridge. In the past, scientists tried to teach robots to notice these mistakes by giving them a checklist of "features" (like "is the bread on the floor?") and asking them to look at data that had already been cleaned up and organized. But the authors of this paper say, "Wait a minute! That's like trying to learn to ride a bike by looking at a diagram of a bike in a perfectly lit studio. Real life is messy!"

This paper introduces ERR@HRI 3.0, a challenge designed to help robots get better at spotting their own blunders—and even guessing them before they happen—by looking at real, messy human reactions.

The Two Big Missions

The challenge gave researchers two different "video game levels" to solve, using raw, unedited webcam footage of people watching robots (and humans) fail.

Level 1: The "Oops!" Detector (The BAD Dataset)
Imagine 45 people sitting in front of a screen, watching videos where a robot or a human makes a mistake. The goal here is reactive: Can a computer look at the person's face after the mistake happens and say, "Oh, they just saw a failure!"?

  • The Catch: The videos are raw. The lighting changes, the camera angles are wonky, and the people are sitting in different spots. It's the real world, not a lab.
  • The Data: There were 46 different failure scenarios and 1,645 video clips in total. The clips are about 15.5 seconds long.
  • The Result: The challenge had 3 teams submit models. All of them did better than the paper's own "baseline" model (which was basically a standard neural network trying its best). The baseline model scored a 0.502 macro F1 (a specific score for how well it balanced detecting both failures and non-failures, rather than just a simple accuracy percentage), but the new models managed to beat that.

Level 2: The "Wait, That's a Bad Idea!" Predictor (The Bad Idea Dataset)
This is the cooler, sneakier part. Imagine 29 people watching a video of a robot starting to do something, but the video cuts off before we see if it succeeds or fails. The people have to guess: "Is this going to end well or poorly?"

  • The Goal: Can a computer look at the person's face while they are guessing and tell what they think is going to happen? This is anticipatory. It's about catching the "uh-oh" look before the crash.
  • The Data: There were 30 scenarios and 865 clips. These clips are super short, only about 1.95 seconds long.
  • The Result: Again, the 3 teams who played this level beat the baseline. The baseline model had an AUC-ROC score of 0.564 (a measure of how well the model can distinguish between a "good" and "bad" guess, where 0.5 is random guessing), showing that predicting the future based on a split-second facial expression is really, really hard.

What This Paper Says "No" To

The authors are very clear about what doesn't work well for this problem, and they explicitly ruled out the old ways of doing things:

  1. No "Pre-Cleaned" Data: The paper notes that most prior work relied on pre-extracted features, which limited the types of methods researchers could use. To fix this, the challenge released raw video data. This design choice wasn't to ban feature engineering, but to enable researchers to apply modern, end-to-end learning methods directly on pixel data and see if they could handle the messiness of real life better than the old approaches.
  2. No "Lab-Perfect" Settings: They rejected the idea that robots should only learn in controlled rooms with perfect lighting. They wanted the messiness of real crowdsourced data (different cameras, weird angles) because that's what robots will actually face in the real world.
  3. No "Just Reacting": They argued that waiting for a mistake to happen and then fixing it is too late. They pushed for anticipation—figuring out something is wrong before it fully happens.

How Sure Are They?

The paper doesn't claim they've solved the problem of robot safety forever. Instead, they suggest that their new datasets and methods are a better starting point.

  • They measured the results using specific math scores (like macro F1 and AUC-ROC) on a test set that the teams had never seen before.
  • They found that while the new models were better than the old baselines, the "anticipation" task (Level 2) is still tricky. The baseline took 35.6% of the clip time to even start detecting the signal, compared to 8.8% for the reactive task. This suggests that reading a "bad idea" face is much harder than reading an "oops" face.
  • The authors state that three teams submitted valid models, and all of them surpassed the baselines. They don't say the models are perfect, just that they are a step forward.

The Takeaway

Think of this paper as the organizers of a science fair who said, "Stop building robots that only work in a glass box. Let's see if they can handle the chaos of a real living room." They provided two new sets of messy, real-world videos and challenged the smartest coders to build models that can spot a robot's mistake or guess a disaster before it happens. The results show that while it's possible to do better than the old methods, the "crystal ball" ability to predict errors is still a work in progress. The authors hope these tools will help build robots that are more aware, more trustworthy, and ready for the messy reality of human life.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →