← Latest papers
🤖 AI

When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

This paper introduces MAVEN, a multi-agent framework that enhances cultural fidelity in text-to-video generation through specialized prompt refinement, validated by a new benchmark of 243 culturally grounded prompts and 972 videos across Chinese, American, and Romanian contexts.

Original authors: Shuowei Li, Yuming Zhao, Parth Bhalerao, Oana Ignat

Published 2026-07-21
📖 7 min read🧠 Deep dive

Original authors: Shuowei Li, Yuming Zhao, Parth Bhalerao, Oana Ignat

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a movie where the director asks the computer, "Show me a person dancing in a famous city." In the world of artificial intelligence, this is called "text-to-video" generation. For a long time, the main goal was just to make the video look real—like the person actually exists and the lighting looks correct. But there's a hidden problem: computers are often terrible at understanding culture. They might know what a "dance" is, but they don't know the difference between a traditional Chinese fan dance, an American hip-hop routine, or a Romanian folk dance. They might mix them up, or worse, create a weird, generic version that doesn't feel like any of them. This matters because as AI starts making videos for stories, education, and ads, we don't want it to accidentally erase traditions or make up fake stereotypes. We want it to get the details right, especially when a scene mixes people, actions, and places from different parts of the world.

This paper tackles that exact problem by introducing a new way to teach AI how to handle "multicultural" scenes. The researchers built a special test set (a benchmark) with 243 different prompts and 972 videos to see how well current AI models can handle mixing cultures. They found that standard AI models struggle when asked to combine, say, an American person eating Chinese food at a Romanian castle. To fix this, they tried a clever trick called MAVEN (Multi-Agent Video Enrichment for cultural Narrative). Instead of asking one big, general AI brain to do everything, they broke the job down into three smaller, specialized "agents": one who is an expert on people (appearance), one on actions (how things are done), and one on locations (landmarks).

The study suggests that having these three experts work together in parallel (all at the same time) works much better than having one generalist do it all or having them work one after another. When the researchers tested this, the parallel team of experts significantly improved how culturally accurate the videos were, especially for landmarks, without messing up the video quality or making the characters jittery. They also discovered that while standard computer tests (using math to compare images) can tell if a video looks "okay," they often miss the subtle, deep cultural details that a more advanced AI "judge" can spot. The paper concludes that to make AI videos truly respectful and accurate, we need to stop treating culture as a single, blurry blob and start breaking it down into specific, expert-driven parts.

The Core Idea: Why One Brain Isn't Enough

Think of the current AI video generators like a single, very talented but overworked chef. If you ask this chef to make a dish that combines Italian pasta, Japanese sushi, and Mexican tacos, they might try to mash it all into one giant, confusing burrito-pasta-sushi roll. It might look like food, but it won't taste like any of the real dishes. The chef knows what "food" is, but they don't have the specific, deep knowledge of how to make each one authentic.

The authors of this paper argue that culture is "compositional." This means a scene is built from different parts: the person (what they look like), the action (what they are doing), and the location (where they are). A person from one culture might do an action from another culture in a place from a third culture. The paper suggests that trying to get one AI to understand all three at once is too hard. Instead, you need a team.

The Solution: The MAVEN Team

The researchers created a system called MAVEN. Instead of one chef, they hired a team of three specialist sous-chefs:

  1. The Person Agent: An expert who knows exactly how people from a specific culture look (hair, clothes, features).
  2. The Action Agent: An expert who knows the specific details of how a cultural action is performed (how to hold a musical instrument, how to dance).
  3. The Location Agent: An expert who knows the visual details of famous landmarks (the specific architecture, the lighting, the surroundings).

They tested four different ways to run this team:

  • The Solo Chef (Base): No help at all. Just the raw prompt.
  • The General Manager (Single-Agent): One agent tries to fix the person, action, and location all at once.
  • The Assembly Line (Sequential): The agents work one after another. The Person agent fixes the prompt, then passes it to the Action agent, who passes it to the Location agent.
  • The Parallel Team (MAP): All three agents work at the same time, each fixing their own part, and then a "Fuse Agent" combines their work into one perfect prompt.

What They Found: The Power of Working Together

The results were clear: The Parallel Team (MAP) won.

When they tested these methods on 972 generated videos (spanning Chinese, American, and Romanian cultures), the Parallel Team produced the most culturally accurate videos.

  • The Score: The Parallel Team improved the "Cultural Relevance Score" by 4.6% compared to doing nothing, and by 1.6% compared to the Single-Agent approach.
  • The Big Win: The biggest improvement was in locations. The Parallel Team improved the accuracy of landmarks by over 12% compared to the baseline. This makes sense because a single agent might get distracted by the person or the action, but a dedicated Location agent focuses purely on getting the castle or monument right.
  • Cross-Cultural Magic: The system was even more helpful when the prompt was a mix of cultures (e.g., an American person eating Chinese dumplings at a Romanian castle). These "cross-cultural" scenes are usually the hardest for AI, but the Parallel Team narrowed the gap significantly, suggesting that breaking the problem down helps the AI handle complex mixes.

Interestingly, the researchers found that while the Parallel Team made the videos culturally better, it didn't ruin the video quality. The videos still looked smooth and consistent, proving that you can add deep cultural detail without making the AI "hallucinate" or break the video.

The Measurement Problem: Computers vs. Judges

One of the paper's most interesting findings is about how we measure success.

  • The Math Test (CLIP): The researchers used a standard computer program (CLIP) to check if the video matched the text. This program is good at spotting big things but often misses the small, subtle cultural details.
  • The AI Judge (VLM): They also used a more advanced AI (Gemini 2.5 Pro) to act as a human-like judge, looking at the video and reasoning about whether it felt culturally correct.

They found that the Math Test often underestimated the improvements made by the Parallel Team. The AI Judge, however, gave much higher scores to the Parallel Team's videos, noticing the subtle details like the specific way a person held a musical instrument or the exact style of a building. This suggests that our current tools for testing AI might be too simple to catch the real value of cultural refinement.

What This Means (and What It Doesn't)

The paper suggests that to make AI videos that respect and accurately represent the world's diversity, we need to stop treating culture as a single, uniform thing. Instead, we should decompose it into specific parts (people, actions, places) and give specialized experts the job of refining each part.

However, the authors are careful to note the limits of their work. They only tested three cultures (Chinese, American, Romanian) and three types of actions (food, music, dance). They don't claim this solves the problem for every culture in the world. They also admit that their AI model sometimes struggles to show faces clearly (often showing people from the back), which makes it hard to judge cultural details about appearance.

But the core message is a hopeful one: by organizing AI into a team of specialists rather than a single generalist, we can take a significant step toward generating videos that don't just look real, but feel culturally true.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →