VideoNorms: Benchmarking Cultural Awareness of Video Language Models
This paper introduces VideoNorms, a dataset of over 3,000 cultural norm annotations from US and Chinese TV shows created through human-AI collaboration, which reveals that current VideoLLMs struggle more with Chinese cultural contexts and non-verbal evidence, highlighting the need for culturally grounded training and evaluation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to understand human behavior. You show it a video of two people meeting, and you ask, "Are they being polite?" In the world of Artificial Intelligence, this is a huge challenge. We have built "Video Large Language Models" (VideoLLMs)—super-smart computers that can watch videos and talk about what they see. But most of these robots have only been trained on data from one specific corner of the world, mostly the United States. They are great at spotting a cat or a car, but they often stumble when it comes to the invisible rules of society: the unwritten scripts we follow when we shake hands, apologize, or say goodbye. Just like a tourist who doesn't know it's rude to point with their left hand in some countries, these AI models might think a behavior is normal when, in a different culture, it's actually a major faux pas. This paper asks a critical question: Can these video-watching robots understand the cultural "vibe" of different societies, or are they just guessing based on what they've seen in American movies?
To answer this, a team of researchers created a new test called VIDEONORMS. Think of this dataset as a giant, global "spot the difference" game, but instead of finding hidden objects, the AI has to spot whether people are following or breaking social rules. The researchers didn't just ask the AI to guess; they built a "Human-AI Team." First, a super-powerful AI (called Gemini) watched thousands of 15-second clips from popular TV shows in the US and China and wrote down what it thought the social rules were. Then, real human experts—people who grew up in those specific cultures—stepped in to act as editors. They checked the AI's work, fixing mistakes and adding details about body language and tone that the robot might have missed. This process resulted in a massive library of over 3,000 human-verified judgments about whether characters in TV shows were being polite or rude.
When the researchers tested seven different open-source video AI models on this new dataset, they found some surprising gaps in the robots' cultural intelligence. First, the models were much better at understanding American social norms than Chinese ones. It's as if the robots had studied hard for a test on US culture but barely opened their textbooks for China. Specifically, the models struggled to predict when a Chinese character was following a rule, often getting confused about what "polite" looked like in that context. Second, the robots had a much harder time explaining why something was rude or polite if the clue was non-verbal. If a character crossed their arms or avoided eye contact, the AI often missed the signal, whereas it was much better at catching verbal clues like the words "sorry" or "thank you."
The study also tested whether simply making the AI "bigger" or smarter would fix these problems. They tried using larger models and feeding them more video frames, but the results didn't improve much. This suggests that the problem isn't just about how much data the robot has memorized; it's about what kind of data it has. The researchers also proved that you can't just read the script (the words spoken) to understand the scene; you actually need to watch the video. The visual cues are essential for getting the cultural context right.
In short, the paper reveals that while our video-watching AI is getting better at seeing the world, it still has a long way to go before it can truly understand the diverse ways humans behave across different cultures. The robots are currently like tourists who know the dictionary but haven't learned the local customs, and simply making them bigger won't teach them the rules of the road. To build truly culturally aware AI, we need to train them on a much wider variety of human stories, not just the ones from Hollywood.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.