Multimodal Rapport Estimation in Real-World HRI
This study demonstrates that multimodal fusion of zero-shot LLMs (specifically Gemini 2.5 Flash) with pretrained audio and visual models (HuBERT and V-JEPA) effectively estimates third-party-rated rapport in real-world HRI settings, outperforming individual models and highlighting the necessity of accounting for contextual variability like interaction duration and group size.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the bustling world of human-robot interaction, researchers have long sought a way to measure the invisible thread that connects a person to a machine. This connection, often called rapport, is not a static personality trait but a dynamic feeling that emerges when two parties pay attention to one another, coordinate their actions, and share a sense of warmth. While scientists have successfully studied this bond in quiet, controlled laboratories where every variable is managed, the real world is far messier. In a busy store or a public square, people wander in and out of conversations at their own pace, groups of friends might gather around a robot simultaneously, and no one is forced to stay. The central question for modern robotics is whether the tools used to measure connection in the lab can still work when the environment is uncontrolled and the participants are free to leave whenever they wish.
To answer this, a team of researchers set up a social robot in a Japanese drugstore, a setting where customers could approach freely without any prior instruction. Over six days, the robot, a small tabletop figure named Sota, engaged in casual conversations with visitors while a human operator guided its speech and gestures from a distance. The team recorded 62 of these interactions, capturing video and audio of people who ranged from solitary shoppers to small groups of friends. They then asked three independent human judges to watch the recordings and rate the quality of the relationship between each person and the robot using a standardized scale. This scale measured how attentive and coordinated the interaction felt, providing a reliable score of the rapport achieved in each moment.
With this real-world data in hand, the researchers tested whether computers could learn to predict these human ratings automatically. They compared two different approaches. The first approach used specialized computer models trained on massive amounts of text, audio, and video data to extract specific features from the recordings. The second approach utilized powerful, general-purpose artificial intelligence systems known as large language models. These systems were given the conversation transcripts and, in some cases, the audio and video files, and asked to act as a third-party judge to rate the rapport themselves, just as the human experts had done.
The results revealed a surprising strength in the general-purpose artificial intelligence. When provided with just the text of the conversation, one of these large models performed exceptionally well, capturing the nuances of the interaction better than the specialized models that relied on audio or visual data alone. The text-only version of this model was able to predict the rapport scores with a high degree of accuracy, suggesting that the words people use and the flow of their conversation carry the most significant clues about their connection to the robot. However, the specialized models were not useless. The researchers found that while the text-based model was strong, it missed certain details that the audio and visual models could catch. When the researchers combined the predictions of the text-based model with those of the audio and visual models, the result was the most accurate system of all. This combined approach outperformed any single method, indicating that the different types of data offer complementary insights rather than competing ones.
The study also highlighted how the length of an interaction and the number of people involved affected the ability to measure rapport. In the controlled environment of a laboratory, interactions are often long and involve just two people. In the drugstore, interactions were brief, averaging just over 54 seconds, and often involved multiple people joining in. The researchers found that the specialized models struggled more as the number of participants increased, likely because the complex dynamics of a group conversation made it harder to isolate individual signals. In contrast, the general-purpose artificial intelligence model remained robust, maintaining its high accuracy even when three people were talking to the robot at once. This suggests that these advanced systems are better equipped to handle the chaotic, multi-person nature of real-world encounters than the traditional methods developed for the lab.
Ultimately, the work demonstrates that estimating the quality of human-robot relationships in the real world is possible, but it requires a different strategy than what has been used in the past. The findings suggest that relying solely on specialized models trained for specific tasks may not be enough when users are free to engage and disengage as they please. Instead, a hybrid approach that leverages the broad understanding of general artificial intelligence while supplementing it with specific audio and visual cues offers the most reliable path forward. This insight is crucial for the future of social robots, as it points toward systems that can truly adapt to the unpredictable and varied nature of human life in public spaces.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.