← Latest papers
💬 NLP

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

The paper introduces FriendBench, a benchmark for inferring dyadic familiarity from 20-second ice-breaker clips, revealing that while top multimodal models match human accuracy, they rely on different cues and exhibit a bias toward classifying pairs as strangers, whereas humans uniquely leverage visible behavioral cues beyond speech.

Original authors: Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino, Antony D'Avirro, Benjamin Peloquin

Published 2026-08-03
📖 3 min read☕ Coffee break read

Original authors: Jeffrey M. Girard, Jason Z. Zheng, Jacqueline R. Vertino, Antony D'Avirro, Benjamin Peloquin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking into a room where two people are talking. You don't know their names, and you haven't heard what they are saying. Yet, within seconds, you can often tell if they are best friends catching up or two strangers awkwardly breaking the ice. This superpower is called social intelligence. It's the ability to read the "vibe" of a situation by watching how people move, pause, and look at each other, rather than just listening to their words. For decades, scientists have studied how humans do this, but now, we are asking a big question: Can computers do it too? As we start building AI companions, meeting assistants, and care robots, these machines need to understand human relationships just as well as we do. But to teach them, we first need a fair test to see if they are actually "reading" the room or just guessing.

Enter FriendBench, a new experiment designed to test exactly this. The researchers set up a digital "ice-breaker" challenge. They took 96 pairs of people and recorded 20-second clips of them answering the same silly question: "Would you rather fly or be invisible?" Some pairs were already friends or family (familiar), while others were meeting for the first time (strangers). Because everyone answered the same question, the words they used didn't give away the answer. The only clues were in how they talked and acted.

The researchers then asked two groups to solve the mystery: a large crowd of human volunteers and 26 different AI models from seven major tech companies. They tested the AI and the humans in three ways: by reading only the text transcript, by listening to the audio, and by watching the video.

Here is what they found. First, the text-only version was a dead end for everyone. Without seeing or hearing the people, neither humans nor AI could do much better than random guessing. This makes sense because the conversation topic was controlled to hide the relationship.

When they added sound, things got better. Both humans and the best AI models improved, showing that tone of voice and pauses carry real clues. But the real magic happened when they added video. Humans got significantly smarter when they could see the people's faces and body language. They used those visual cues to boost their accuracy even further. The AI models, however, hit a wall. While the best AI models got just as accurate as the human crowd in the video clips, they didn't get there the same way.

The AI models didn't actually "see" the friendship in the way humans did. Instead, they leaned heavily on a pattern: they guessed "stranger" almost all the time. Because there were equal numbers of friends and strangers in the test, this bias accidentally kept their overall score high, but it meant they were failing to spot the friends. Humans, on the other hand, stayed balanced, correctly identifying both friends and strangers.

In short, the best AI models can now match human accuracy on this task, but they are doing it by relying on a strategy that assumes everyone is a stranger, rather than by truly understanding the social dance. They hear the voice, but they aren't quite reading the body language yet. This suggests that while AI is getting good at social tasks, it still has a long way to go before it can truly understand human relationships the way we do.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →