← Latest papers
💻 computer science

HUI360: A 360{\deg} Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation

This paper introduces HUI360, the largest in-the-wild 360-degree egocentric dataset for human-robot interaction anticipation, featuring over 1 million pre-processed annotations, a pipeline for automatic interaction labeling, and benchmark baselines to advance the generalization of proactive robotic behaviors.

Original authors: Raphael Lorenzo-Louis, Fabio Amadio, Bertrand Luvison, Serena Ivaldi

Published 2026-08-12
📖 4 min read☕ Coffee break read

Original authors: Raphael Lorenzo-Louis, Fabio Amadio, Bertrand Luvison, Serena Ivaldi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking through a busy park, and a friendly robot rolls up to you. Does it stop and wait for you to say hello, or does it guess you're coming and wave before you even get close? This is the heart of a field called Human-Robot Interaction (HRI). It's the science of teaching machines to understand us, not just as moving objects, but as people with intentions. A key part of this is "anticipation"—the ability to predict what a human will do next. Think of it like a dance partner who knows the next step before the music even changes. If a robot can anticipate that you want to chat, it can be more helpful and less awkward. But for a robot to learn this dance, it needs to practice. And to practice well, it needs a massive library of real-life examples, not just rehearsed scenes in a studio.

This is exactly what the researchers behind the paper "HUI360" set out to build. They created a giant, open-source library of video data called HUI360, which is currently the largest dataset of its kind for teaching robots how to guess when humans want to interact with them. Instead of using actors in a controlled lab, they sent a mobile robot out into the "wild"—nine different real-world locations like cafeterias, hallways, and public squares—over several months. The robot, wearing a 360-degree camera like a fish-eye lens, watched thousands of people walk by. Some ignored it; others stopped to pick up stickers, drop trash, or grab food. The team used smart computer vision tools to automatically track every person, map their body movements, and flag the exact moment they physically touched the robot. They even cleaned up the data by hand to fix any mistakes, ensuring the "lessons" the robot learns are accurate.

The paper also introduces a set of "baselines," which are like the robot's first practice tests. They trained simple computer brains (using methods like Random Forests and LSTMs) to look at a person's movement and guess if they were about to interact. The results suggest that these models can get pretty good at the job, especially when they use detailed body maps (like facial and hand keypoints) to make their guesses. However, the paper also highlights a tricky reality: when the robot is tested in a completely new environment it hasn't seen before, or when the robot itself changes (like swapping a trash-can robot for a service robot), the prediction gets a bit fuzzier. It's like a student who aced a math test in one classroom but stumbled when the teacher moved to a different room. The authors show that while current methods work, they need to get much better at adapting to new places and new robot bodies to be truly useful in the real world.

Crucially, the paper draws a hard line in the sand about what counts as an "interaction." The researchers decided to ignore the tricky stuff, like when someone just stares at the robot out of curiosity or smiles at it. Those moments are too subjective and hard to agree on. Instead, they focused only on physical interactions—when a person actually touches the robot, throws something in its bin, or picks up an object from it. This makes the data objective and easy to verify. They argue that while we might wish robots could read minds or understand subtle glances, starting with clear, physical actions is the only way to build a reliable foundation for learning.

So, what did they find? The HUI360 dataset, containing over 1 million annotations of people and their movements, proves that robots can learn to anticipate physical interactions with surprising accuracy. The models they tested could predict an interaction a second or two before it happened, which is enough time for a robot to turn its head or prepare a greeting. But the paper also suggests that this is just the beginning. The models struggle when the environment changes drastically, indicating that future robots will need to be much more flexible to handle the chaos of the real world. By releasing this massive dataset and the tools to create it, the authors are handing the keys to the rest of the scientific community, inviting everyone to build better, more socially aware robots that can finally learn to dance with us in the wild.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →