CULTURESCORE: Evaluating Cultural Faithfulness in Video Generation Models
This paper introduces CultureScore, a novel evaluation framework that decomposes cultural faithfulness into identity, context, and behavior dimensions to reveal that current state-of-the-art video generation models significantly struggle with culturally accurate representation, often prioritizing visual quality over cultural correctness.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a magical movie camera that can create any scene you describe. You tell it, "Show me a family in India celebrating a festival," and it spits out a beautiful, high-definition video. But here's the catch: the family in the video is wearing Western suits, eating at a fancy dining table, and shaking hands instead of bowing with folded hands.
To a standard quality checker, this video might get a perfect score because the lighting is great, the people look real, and the camera moves smoothly. But to a local Indian person, the video feels wrong, like a costume party where everyone got the costumes mixed up.
This paper introduces a new way to test these "magic movie cameras" (AI video generators) called CULTURESCORE. It argues that just because a video looks pretty doesn't mean it looks true to a specific culture.
Here is the breakdown of their findings using simple analogies:
1. The Problem: The "Perfectly Wrong" Video
Current tools that grade AI videos (like VideoScore) are like art critics who only look at the frame. They check: "Is the picture blurry? Is the color bright? Does the person move smoothly?"
- The Flaw: If the AI replaces a traditional Indian greeting (a Namaste) with a Western handshake, the art critic gives it a high score because the handshake is drawn perfectly.
- The Reality: For someone from that culture, the video is a failure. It's like serving a delicious-looking cake that tastes like soap.
2. The Solution: The "Three-Layer Cake" (CULTURESCORE)
The authors propose a new grading system called CULTURESCORE. Instead of just looking at the whole picture, they slice the video into three specific layers to see where the AI messed up:
- Layer 1: Identity (Who is in the room?)
- The Analogy: Are the actors wearing the right clothes? Do they look like the people they are supposed to be?
- Example: If the prompt says "Japanese colleagues," are they dressed in typical Japanese business attire, or do they look like generic Western office workers?
- Layer 2: Behavior (What are they doing?)
- The Analogy: Are they moving with the right rhythm?
- Example: A Namaste isn't just hands together; it's a specific bow and timing. A handshake is a different rhythm. The AI often gets the "dance steps" wrong, even if the actors look okay.
- Layer 3: Context (Where are they?)
- The Analogy: Is the background set up correctly?
- Example: If it's a family dinner in a specific culture, are they sitting on the floor around a low table (a dastarkhwan) or at a tall Western dining table?
3. The Big Discovery: The "Inverted World"
The researchers tested three of the smartest video AI models (Veo, LTX-2, and Wan) with prompts from 10 different countries. They found a shocking pattern:
- The "Hollywood" Model: One model (LTX-2) made the most "cinematic," high-quality videos. It got the highest scores on traditional quality checks.
- The "Cultural" Model: Another model (Wan 2.2) made videos that were less "pretty" but much more culturally accurate.
- The Twist: When real humans from those countries watched the videos, they hated the "Hollywood" model and loved the "Cultural" model.
- The Metaphor: It's like a restaurant where the most expensive, beautifully plated dish tastes terrible to the locals, while the simple, home-cooked meal is the favorite. The standard judges (VideoScore) were praising the wrong dish.
4. The "Magic Words" Test
The researchers also tested if the AI actually understands culture or if it just memorizes keywords.
- They asked the AI: "Show me a Muslim Indian family eating." (AI gets it right).
- Then they asked: "Show me a Muslim family eating." (AI gets it wrong).
- The Lesson: The AI isn't truly understanding the culture; it's just looking for the word "India" or "Japan" as a trigger. If you take away the country name, the AI forgets the cultural rules. It's like a student who memorized the answer key but doesn't understand the math.
5. The Hardest Part: The "Dance"
Out of the three layers, Behavior (the actions and gestures) was the hardest for the AI to get right.
- Even when the researchers gave the AI very detailed instructions, the AI still struggled to get the movement right.
- The Analogy: The AI can draw a person wearing a kimono perfectly (Identity) and put them in a Japanese garden (Context), but when it tries to make them bow, it looks stiff or robotic. The "dance" of culture is still very hard for computers to learn.
Summary
The paper concludes that we cannot just trust the "pretty" scores when judging AI videos. A video can be technically perfect but culturally offensive or inaccurate. To make AI fair and useful for the whole world, we need to check the Who, the What, and the Where separately, and we need to listen to the people who actually live in those cultures, not just the machines grading the picture quality.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.