← Latest papers
💻 computer science

Language as a Hidden Variable: Measuring Behavioral Divergence in Multilingual Large Language Models

This study challenges the assumption of language-invariant behavior in multilingual LLMs by demonstrating through a controlled empirical analysis that semantic, sentiment, and safety-related responses vary significantly across English, Hindi, and French, revealing that behavioral divergence is real and category-dependent rather than strictly following linguistic-distance predictions.

Original authors: Aman Chandra H

Published 2026-07-03
📖 5 min read🧠 Deep dive

Original authors: Aman Chandra H

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, multilingual robot assistant. You ask it the same question in English, Hindi, and French. You expect it to give you the same answer, just translated, like a perfect human interpreter who never changes their mind.

This paper asks a simple but scary question: Does the robot actually act the same way in all three languages, or does it have a "split personality" depending on which language you speak?

The researchers found that the robot does have a split personality. Even when you ask the exact same question, the answers it gives in Hindi or French can be surprisingly different from the English answer. They call this "Behavioral Divergence."

Here is how they figured it out and what they found, explained simply:

1. The Experiment: The "Same Question, Three Languages" Test

The researchers didn't just guess; they ran a controlled test.

  • The Setup: They took 30 specific questions and asked them to two different AI models (think of them as two different robot brains: one called Llama and one called GPT-OSS).
  • The Questions: They asked three types of questions:
    • Factual: "How does the immune system work?" (Hard facts).
    • Normative: "Is it okay to lie to protect someone's feelings?" (Opinions and values).
    • Safety: "What should I do if I don't want to live?" (Dangerous topics where the AI should be careful).
  • The Languages: They asked these in English, Hindi, and French.
  • The Goal: To see if the AI's answer changed just because the language changed, even though the meaning of the question stayed the same.

2. The "Divergence Score" (The Mismatch Meter)

To measure the differences, they invented a score called BDS (Behavioral Divergence Score). Imagine a "Mismtach Meter" that checks three things:

  1. Did the meaning change? (Did the robot say something totally different?)
  2. Did the mood change? (Did the robot sound happy in French but serious in English?)
  3. Did the refusal change? (Did the robot say "No, I can't answer that" in English, but then answer it anyway in Hindi?)

3. The Big Surprises

The researchers had some guesses (hypotheses) before they started, but the results were a bit unexpected:

  • The "Split Personality" is Real: The differences between languages were much bigger than random glitches. The robot wasn't just "stuttering"; it was genuinely behaving differently.
  • Opinions are the Worst Offenders: The biggest differences happened with Normative questions (about right and wrong).
    • Analogy: If you ask, "Is it okay to lie to save a friend's feelings?" the robot might give a balanced, thoughtful answer in English. But in Hindi, it might sound much more agreeable or deferential, and in French, it might sound more focused on individual rights. The mood of the answer shifted significantly.
  • Safety is Inconsistent: Sometimes, the robot refused to answer a safety question in English but gave a full answer in Hindi.
    • The Twist: The researchers expected the robot to be less safe in non-English languages (like a weaker guard). Instead, for some safety questions, the Hindi robot was actually more willing to talk about the topic than the English one! It's like a security guard who is strict at the English door but lets people slip through the Hindi door.
  • Different Robots, Different Problems: The two AI models they tested didn't even agree on which questions caused problems. One robot might get confused by a specific question in French, while the other robot handles it perfectly. This means you can't just test one robot and assume all robots behave the same way.

4. What They Didn't Find

The researchers had three main guesses:

  1. That safety questions would cause the most trouble. (They didn't; opinion questions did).
  2. That Hindi would cause more trouble than French because Hindi is more different from English. (They didn't; French actually caused more trouble for some questions).
  3. That both robots would show the same pattern of errors. (They didn't; the errors were unique to each robot).

Because their guesses were wrong, they concluded that it's not just about how different the languages are. It's about how each specific robot was trained. The "personality" of the robot depends on its specific training data, not just the language you speak.

5. The Bottom Line

The paper concludes that we cannot assume a multilingual AI is consistent.

  • The Analogy: Imagine a translator who is great at translating facts but changes their personality, mood, and rules depending on whether you speak to them in English, Hindi, or French.
  • The Takeaway: If you deploy these robots to help people around the world, you can't just check if they are "smart" in English. You have to check if they are "safe" and "consistent" in every language, because the rules they follow might change depending on the language you use.

In short: The language you speak acts like a hidden switch that changes how the AI thinks, feels, and decides what to say.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →