TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories
This paper introduces TALES, a comprehensive framework comprising a taxonomy of cultural misrepresentations and a large-scale evaluation of 6 LLMs, which reveals that 88% of AI-generated stories about Indian cultures contain errors, particularly affecting mid- and low-resourced languages and peri-urban narratives.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, well-read robot friend who loves to tell stories. You ask it to write a tale about a wedding in a small village in India, or a childhood memory in a bustling city. You expect the robot to capture the local flavor—the specific snacks, the right family titles, the unique traditions.
But according to this paper, TALES, that robot friend is often getting the details wrong. In fact, it's like asking a tourist who has only read a travel brochure to describe your hometown; they might get the big landmarks right, but they'll likely mix up the local slang, the food, and the daily routines.
Here is a simple breakdown of what the researchers found, using some everyday analogies:
1. The Problem: The Robot's "Travel Brochure" Brain
The researchers wanted to know: When AI tells stories about Indian culture, does it get it right?
They found that 88% of the stories generated by popular AI models contained mistakes. These weren't just tiny typos; they were cultural blunders.
Think of it like a chef trying to cook a regional dish.
- The Mistake: The chef puts a dessert ingredient (like sugar) into a savory dish (like a curry) because they think "Indian food" is all one thing.
- The Reality: The researchers found the AI often confused specific regional foods, used the wrong family titles (calling a male relative "aunt"), or described events that simply don't happen in real life (like a school bus having a ticket conductor).
2. The Tool: The "Cultural Cheat Sheet" (TALES-Tax)
Before they could count the mistakes, the researchers needed a way to categorize them. They didn't just guess; they asked real people from India (9 focus groups and 15 individual surveys) to read the AI stories and point out what felt "off."
Based on this feedback, they created a Cultural Cheat Sheet (called TALES-Tax) with seven categories of errors:
- Cultural Inaccuracy: Getting facts wrong (e.g., saying a specific snack is eaten for breakfast when it's actually a snack).
- Unlikely Scenarios: Making up events that feel fake (e.g., a street performer playing a complex classical instrument during a morning school assembly).
- Clichés: Relying on stereotypes (e.g., assuming everyone in a region drinks a specific type of coffee or wears traditional festival clothes to school).
- Oversimplification: Smoothing over unique details (e.g., calling all Indian art "Rangoli" instead of using the specific local name).
- Factual Errors: Getting geography or history wrong (e.g., saying there is a desert next to a city that is actually green).
- Linguistic Errors: Spelling names wrong or using family words incorrectly.
- Logical Errors: Breaking the rules of how things work (e.g., a character bathing in a drinking water well).
3. The Big Test: The "Taste Test" with 108 Experts
The researchers didn't just look at one story. They hired 108 experts from 71 different regions across India. These were native speakers who knew their local cultures inside and out.
They asked these experts to read 540 stories written by 6 different AI models (including big names like GPT-4 and Gemini) in 14 different languages.
The Results:
- The "Rich" vs. "Poor" Gap: The AI was much better at telling stories about big, famous cities (Tier-1) and in English. When the story was set in a smaller town or written in a less common Indian language, the mistakes quadrupled.
- The "Generic" Trap: Some AI models tried to avoid mistakes by writing very boring, generic stories with no cultural details at all. Others tried to be creative but got the details wrong. Neither approach was good.
- Relatability: Because of these mistakes, the stories felt "unrelatable" to the people they were supposed to represent. It's like watching a movie about your hometown where the actors speak with a fake accent and eat the wrong food; you just can't connect with it.
4. The Big Surprise: The "Knowledge vs. Performance" Paradox
Here is the most interesting part. The researchers wondered: "Is the AI just ignorant? Does it not know the facts?"
To test this, they turned the mistakes into a quiz (called TALES-QA). They asked the AI direct questions like, "What is the traditional food served at this wedding?" or "Where is this statue located?"
The Shocking Result:
- When asked as a quiz, the AI got the answers right 76% of the time (in English) and 60% of the time (in Indian languages).
- The Conclusion: The AI knows the facts. It has the information in its brain. But when it tries to write a story, it fails to use that knowledge correctly.
The Analogy:
Imagine a student who can ace a multiple-choice test on history (getting 90% right) but, when asked to write a history essay, makes up wild, incorrect facts. The student isn't "ignorant"; they just can't apply what they know in a creative, open-ended situation.
Summary
The paper concludes that current AI models are like overconfident tourists. They have read the guidebooks (they have the data), but when they try to tell a story about a local culture, they rely on stereotypes, mix up details, and fail to capture the true "flavor" of the place. This is especially true for smaller towns and less common languages.
The researchers hope their "Cultural Cheat Sheet" and "Quiz" will help developers fix this, so that future AI can tell stories that feel real and respectful to the people they are about.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.