← Latest papers
🧬 biology

Behavioral Geometric Supervision Aligns Video Foundation Models with Human Social Perception

This paper introduces Behavioral Geometric Supervision (BGS), a hybrid objective that aligns video foundation models with human social perception by constraining embedding geometry to match human similarity judgments, enabling models to significantly outperform language-based baselines, develop interpretable social-affective attributes, and shift attention to socially informative regions without explicit training on those features.

Original authors: Kathy Garcia, Leyla Isik

Published 2026-05-14
📖 4 min read☕ Coffee break read

Original authors: Kathy Garcia, Leyla Isik

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

The Big Problem: AI Watches, But Doesn't "Get" People

Imagine you show a video of two people arguing to a very smart AI and a human.

  • The Human instantly understands the social nuance: "They are frustrated but still caring for each other," or "This is a playful competition."
  • The AI (specifically the best video models available today) mostly sees pixels moving. It might recognize that "a person is raising a hand" or "someone is shouting," but it fails to understand the relationship or the feeling between the people.

The researchers found that if you asked an AI to guess which two of three video clips were most similar, it performed worse than a simple text-based AI that just read the captions describing the videos. In other words, the AI was better at reading the "movie summary" than actually "watching the movie" to understand the social dynamics.

The Solution: Teaching AI with "Odd-One-Out" Games

The researchers wanted to fix this without expensive brain scans or massive amounts of new data. They introduced a method called Behavioral Geometric Supervision (BGS).

Think of it like teaching a child to recognize emotions not by giving them a textbook, but by playing a game:

  1. The Game: Show the AI three short video clips. Ask it: "Which one is the odd one out?" (e.g., "Two clips show friends hugging, one shows strangers fighting. Which is different?")
  2. The Human Teacher: Humans played this game thousands of times, creating a massive map of how we see social similarities.
  3. The Lesson: The researchers didn't just tell the AI the right answer. They adjusted the AI's internal "brain" so that its mathematical understanding of "distance" between videos matched the human map. If humans thought Video A and Video B were very similar, the AI was forced to make its internal representation of A and B very close together.

The Magic Ingredients

To make this work efficiently, they used two main tools:

  • LoRA (Low-Rank Adaptation): Imagine the video AI is a giant, heavy library. Instead of rebuilding the whole library, they added a small, lightweight "notebook" (updating less than 2% of the AI's brain) that learned the new social rules. This was fast and cheap.
  • The Hybrid Loss (Local + Global): They taught the AI in two ways at once:
    • Local: "Make sure this specific pair of videos is closer than that one" (like getting the specific game answers right).
    • Global: "Make sure the whole map of relationships looks like the human map" (like understanding the general shape of the social world).

The Results: What Changed?

After this training, the AI didn't just get better at the game; it fundamentally changed how it "saw" the world.

  1. It Beat the Text AI: The trained video models became better at matching human social judgments than the text-based models that read captions. This proves the AI learned something visual that words alone couldn't capture.
  2. It Learned "Invisible" Traits: Without ever being told what "anger," "joy," or "dominance" meant, the AI started to recognize these concepts on its own. It's like a student who, after studying many examples of "friendship," suddenly realizes they can identify "trust" without anyone ever defining the word.
  3. It Looked at the Right Things: Before training, the AI looked at the background or random body parts. After training, its attention shifted to faces, eyes, and hands—the exact places humans look to understand social interactions.
  4. It Didn't Forget How to Dance: Usually, when you teach an AI a new trick, it forgets old ones. But this AI still knew how to recognize standard actions (like "dancing" or "running") just as well as before. It added social understanding without losing its original skills.

The Bottom Line

The paper shows that video AI models already have the raw data to understand human social interactions, but they are "asleep" to it. By using a small amount of human behavioral data (people playing the "odd-one-out" game), the researchers were able to "wake up" this social understanding.

They proved that you don't need to program AI with complex rules about human behavior; you just need to show it how humans compare videos, and the AI will naturally learn to see the world the way we do.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →