LLMs Infer Cultural Context but Fail to Apply It When Responding
This paper introduces the CAPRI dataset to demonstrate that while large language models can infer cultural context and recall relevant conventions, they frequently fail to apply this knowledge to generate culturally adapted responses unless explicitly guided to do so.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, well-read librarian who has read almost every book in the world. This librarian knows a lot about different countries: they know that people in France use Celsius for temperature, that people in the US use Fahrenheit, and that phone numbers look different in Japan than in Brazil.
However, this paper reveals a funny and frustrating flaw in how this librarian works. They know the facts, but they forget to use them when talking to you.
Here is the breakdown of the research using simple analogies:
1. The Test: The "Cultural Detective" Game
The researchers created a game called CAPRI (Cultural and Pragmatic Response Inference). Imagine a chatbot talking to a user. The user drops subtle hints about where they are from, like mentioning a specific French website (Fnac.com) or using a French phone number format.
The chatbot has two jobs:
- Detect the Clue: Figure out, "Oh, this user is probably French."
- Adapt the Answer: When the user asks, "What's the temperature?" (showing a picture of a thermometer), the bot should say "40°C" (French style) instead of "104°F" (American style).
2. The Problem: The "Know-It-All" Who Doesn't Listen
The researchers tested the smartest AI models available. Here is what they found:
- The Detective Part (Job 1): The AI is amazing at this. If you give it even one or two clues, it correctly guesses the user's country almost 100% of the time. It knows the facts.
- The Adaptation Part (Job 2): This is where it fails. Even though the AI knows the user is French, it often ignores that knowledge and gives the answer in American units (Fahrenheit) anyway.
The Analogy: It's like a waiter who knows you are a vegetarian (because you told them you don't eat meat), but when you order a burger, they still bring you a steak because they are "on autopilot" and didn't connect the dots between your identity and your order.
3. The Fix: "Thinking Out Loud"
The researchers tried a trick to fix this. They asked the AI to think out loud before answering. They said, "First, figure out where the user is from. Then, decide what unit to use based on that."
When they did this, the AI got much better. It was like telling the waiter, "Stop and think: Wait, this person is a vegetarian. I need to swap the steak for a salad." Once the AI was forced to pause and reason step-by-step, it finally started giving culturally appropriate answers.
4. The "Default Setting" Bias
The paper also looked at what happens when the AI has no clues at all (no phone numbers, no websites mentioned).
- Country Guess: If the AI has to guess the country with no hints, it usually guesses the United States.
- Unit Guess: But here is the twist: Even though it guesses the US, it often still uses the Metric system (Celsius, kilometers), which the US doesn't use.
It's as if the AI has a "default personality" that is a mix of a US citizen who speaks like a European. It's not truly neutral; it has its own hidden biases based on where the model was built and what data it was trained on.
5. Subjective Things (Time and "Some")
The researchers also tested "squishy" concepts that don't have strict rules, like what time of day it is (is 6 PM "evening" or "afternoon"?) or how many eggs are in a basket ("a few" vs. "some").
- The Good News: As the AI gets more clues, it starts to change its answers to match different cultures.
- The Bad News: Without clues, the AI's "default" answers often lean toward the culture of the company that built the model (e.g., a Chinese model leans toward Chinese interpretations, a US model leans toward US ones).
The Bottom Line
The paper concludes that current AI models are like encyclopedias that haven't learned to be good conversationalists yet. They store cultural facts in one part of their brain and the ability to use them in another, but they don't automatically link the two.
To make them truly helpful, we can't just rely on them to "know" culture; we have to explicitly teach them to pause, think about who they are talking to, and then adjust their answer. Until then, they might know you are French, but they'll still tell you the temperature in Fahrenheit.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.