← Latest papers
💻 computer science

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning

This paper introduces SIV-Bench, a comprehensive video benchmark comprising 2,792 clips and 5,455 question-answer pairs designed to evaluate Multimodal Large Language Models on social interaction tasks, revealing that while current models excel at scene understanding, they struggle with social state reasoning and dynamics prediction due to misalignment with human thought processes.

Original authors: Fanqi Kong, Weiqin Zu, Xinyu Chen, Yaodong Yang, Song-Chun Zhu, Xue Feng

Published 2026-04-28
📖 5 min read🧠 Deep dive

Original authors: Fanqi Kong, Weiqin Zu, Xinyu Chen, Yaodong Yang, Song-Chun Zhu, Xue Feng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to be a good friend. You can show it a picture of a dog, and it can tell you "that's a dog." You can show it a video of a car crash, and it can say "the car hit the tree." But what happens when you show it a video of two people arguing, or a group of friends laughing at a joke? Can the robot understand why they are arguing, or what they are feeling, or predict what they will do next?

This is the problem the paper SIV-Bench is trying to solve.

The Problem: Robots Are "Socially Clueless"

The authors argue that while AI models (the "brains" behind robots and chatbots) are getting very good at seeing and describing the world, they are terrible at understanding human social interactions.

Think of it like this: An AI might be able to describe a video of a birthday party perfectly ("There is a cake, a child is blowing candles, and people are wearing hats"). But if you ask it, "Is the child happy or embarrassed?" or "Why is the uncle looking at the child with a frown?", the AI often gets it wrong. It sees the actions, but it misses the story behind them.

The Solution: A "Driver's License" for Social AI

To fix this, the researchers built a new test called SIV-Bench. Think of this as a rigorous "driver's license" exam, but instead of testing if a car can drive down a street, it tests if an AI can "drive" through a complex social situation.

The test is built on a simple idea: Social relationships are like different types of games.

  • Family/Friends: You share things freely (like sharing a pizza).
  • Boss/Employee: One person gives orders, the other follows (like a coach and a player).
  • Strangers/Business: You trade things based on value (like buying a coffee).

The test uses 2,792 real video clips from the internet (like TikTok and YouTube) featuring 14 different types of relationships (parents, couples, colleagues, etc.). For each video, they created 5,455 questions to see if the AI can pass the test.

The Three Levels of the Exam

The test is divided into three levels, getting harder as you go up:

  1. Level 1: The "What Do You See?" Test (Social Scene Understanding)

    • The Task: "What is the man wearing?" or "What is the woman doing with her hand?"
    • The Result: The AI is pretty good at this. It's like a robot that can read a menu. It sees the words and objects clearly.
  2. Level 2: The "What Are They Thinking?" Test (Social State Reasoning)

    • The Task: "Why is the teacher smiling at the student?" or "Are these two people friends or enemies?"
    • The Result: This is where the AI struggles. It often gets confused. For example, it might see a boss yelling at an employee and think they are just "colleagues" having a loud chat, missing the power dynamic. It's like a robot that sees a storm but doesn't understand that people are scared.
  3. Level 3: The "What Happens Next?" Test (Social Dynamics Prediction)

    • The Task: "If the boss hadn't yelled, what would the employee have done?" or "What will happen next?"
    • The Result: This is the hardest part. The AI has to imagine a different reality or predict the future based on social rules. It often fails here because it doesn't truly "get" human rules.

The Findings: What the Test Revealed

When the researchers ran the test on the smartest AI models available today, they found some surprising things:

  • Big Brains Help, But Don't Solve Everything: The biggest, most powerful AI models did better than the smaller ones, but even the "smartest" ones still failed a lot of the social questions. They are like a student who memorized the textbook but hasn't lived in the real world.
  • The "Relationship" Confusion: The biggest mistake the AI made was mixing up relationships. It often confused a "Boss and Employee" with "Colleagues," or a "Parent and Child" with "Friends." It couldn't tell the difference between a strict teacher and a strict parent.
  • Words Matter: The researchers found that if they gave the AI the audio (what people are saying) or subtitles, it got much better at the hard questions. This is like giving a student a hint sheet; without the words, the AI is often lost in the visual noise.
  • The "Human Gap": Even the best AI models were far behind actual humans. When humans took the test, they got about 74% right. The best AI got about 62% right. The gap shows that AI is still missing a crucial piece of "social intuition" that humans have naturally.

The Conclusion

The paper concludes that while AI is getting better at seeing the world, it is still learning how to understand people. It can describe the scene, but it often misses the emotional and social story.

The authors hope that by releasing this test (SIV-Bench), other scientists will use it to build AI that is not just smart, but also socially intelligent—AI that can truly understand human feelings, relationships, and behaviors, rather than just guessing based on what it sees on the surface.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →