Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
This paper introduces the Singleton Fleiss's metric to quantify cross-lingual cultural inconsistencies in multilingual LLMs and proposes the C-3PO framework, which uses consensus-driven preference optimization to align model outputs with a fixed persona across languages, effectively mitigating the tendency for prompt language to overwrite user-defined cultural identities.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Core Problem: The "Chameleon" Confusion
Imagine you hire a personal assistant who is supposed to act exactly like you. You tell them, "I am British, and I love British history."
- Scenario A: You ask them in English: "Who is the most famous writer studied in school?" They say, "Shakespeare." (Perfect, this matches your identity).
- Scenario B: You ask the exact same question in Spanish: "¿Qué escritor se estudia en la escuela?" They say, "Cervantes."
The Problem: Even though you told them you are British, the moment you switched languages, they forgot who you are and started acting like a Spanish person instead. They let the language of the question override your identity.
The paper calls this Cross-Lingual Cultural Inconsistency (CCI). It's like a chameleon that changes its color not just to match the background, but to match the language you are speaking, even when you've told it to stay a specific color.
The Solution: A New Way to Measure the Mess
Before fixing the problem, the authors needed a way to measure how confused the AI was. Standard tests often fail because if an AI makes up a fake answer (a "hallucination"), it might still look consistent if it makes up the same fake answer in every language.
The Analogy: Imagine a classroom where students are asked to pick a favorite fruit.
- Old Method: If everyone picks "Banana," the teacher thinks they agree. But what if the teacher didn't actually have bananas? The students just guessed the same wrong thing.
- The Paper's New Method (Singleton Fleiss's ): This is a special scoring system that treats every "wrong" or "made-up" answer as a unique, one-of-a-kind error. If the AI says "Banana" in English but "Flying Pizza" in Spanish, the score tanks. If it says "Flying Pizza" in both, the score stays low because it's still a hallucination. This ensures the AI is actually being consistent with reality, not just repeating its own mistakes.
The Fix: C-3PO (The "Group Consensus" Coach)
The authors created a new training method called C-3PO (Cross-lingual Cultural Consistent Preference Optimisation).
How it works (The Analogy):
Imagine you are trying to teach a student to answer a question correctly, but you don't have a textbook with the right answers. Instead, you ask the student the same question in 8 different languages.
- The Poll: You ask the student: "What is the answer?" in English, Spanish, Chinese, etc.
- The Vote: You look at all 8 answers. If 5 out of 8 languages say "Shakespeare," that becomes the "Consensus" (the group's best guess).
- The Correction:
- If the student answered "Shakespeare" in English, you say, "Good job, keep doing that."
- If the student answered "Cervantes" in Spanish (because the language tricked them), you say, "No, that's wrong. The group voted for Shakespeare. You need to align with the group."
- The Training: The AI is then retrained to prefer the "Group Consensus" answer over the "Language-Specific" answer. It learns that no matter what language you speak, the answer should stay the same if the user's identity is fixed.
Key Findings: Who Gets Hurt the Most?
The paper found that this "Chameleon Confusion" isn't equal for everyone.
- The Rich Languages: Languages like English, Spanish, and Chinese (which have huge amounts of data on the internet) are less confused. The AI handles them better.
- The Poor Languages: Languages like Indonesian, Persian, and Greek (which have less data) suffer much more. The AI is much more likely to forget the user's identity and default to stereotypes when speaking these languages.
- Analogy: It's like a student who is great at math in their native language but gets completely lost and starts guessing randomly when asked to do math in a language they haven't practiced as much.
Why Does This Happen? (The "Deep Dive")
The authors looked inside the AI's "brain" (its internal layers) to see when this mistake happens.
The Discovery:
They found that in the early stages of thinking, the AI is just processing the words. But as it gets closer to giving the final answer (in the later layers), it suddenly "locks in" on the cultural stereotype associated with the language.
- Analogy: It's like a traveler who starts a trip with a map of their home country. As they walk further down the road, they start ignoring their map and just following the local signs, eventually forgetting where they started. The AI "forgets" the user's persona right before it speaks.
Summary
The paper argues that when you explicitly tell an AI who you are, it shouldn't let the language you speak change its answer. They built a new tool to measure when the AI fails at this, and they created a training method (C-3PO) that forces the AI to listen to the "majority vote" of its own multilingual answers, effectively teaching it to ignore the language trap and stick to the user's identity.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.