← Latest papers
💬 NLP

Cross-Lingual Response Consistency in Large Language Models: An ILR-Informed Evaluation of Claude Across Six Languages

This paper presents a novel evaluation framework using ILR Skill Level Descriptions to assess Claude's cross-lingual consistency across six languages, revealing significant domain-dependent variations in length, style, and cultural calibration that highlight the necessity of combining automated metrics with expert qualitative judgment for equitable multilingual AI deployment.

Original authors: Camelia Baluta

Published 2026-05-01
📖 5 min read🧠 Deep dive

Original authors: Camelia Baluta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, multilingual robot named Claude. You ask it the exact same question in six different languages: English, French, German, Spanish, Italian, and Romanian. You might expect the robot to give you answers that are identical in quality, tone, and cultural flavor, just translated.

This paper says: No, that's not what happens.

The author, Camelia Baluta, is like a "language detective" who used a special set of rules (called the ILR framework, usually used to grade human language skills) to investigate how Claude behaves. She treated the robot like a student taking a test in six different languages to see if it was playing fair.

Here is what she found, explained with some everyday analogies:

1. The "Lengthy French vs. Concise German" Effect

The Finding: When asked the same question, the French version of the robot wrote answers about 30% longer than the German version.
The Analogy: Imagine asking a tour guide in Paris and a tour guide in Berlin, "Tell me about this building." The Paris guide might tell you a long, winding story with lots of details and history. The Berlin guide might give you the facts, the date, and the architect, then stop. Both are correct, but the amount of information you get depends entirely on which language you speak.

2. The "Creative vs. Technical" Split

The Finding: The robot's answers were most different when asked to be creative (like writing a story) or emotional. They were most similar when asked about technical facts or language rules.
The Analogy: Think of the robot as an actor.

  • Technical Mode: When asked to explain a math problem, the actor puts on a "uniform" and speaks the same script in every language. It's boring but consistent.
  • Creative Mode: When asked to write a poem about rain, the actor takes off the uniform. In Romanian, the rain might feel like a "living city breathing." In English, it might just be "grey clouds." The robot isn't just translating; it's actually feeling the story differently depending on the language it's speaking.

3. The "Ambiguity Dance"

The Finding: When the question was vague (e.g., "They said we shouldn't come. What do we do?"), the robot reacted differently based on cultural habits.
The Analogy: Imagine a group of friends trying to solve a mystery.

  • The German Friend: Says, "I can't answer until you tell me exactly who 'they' are." (They need precision).
  • The Spanish/Italian Friend: Says, "Well, assuming they mean the neighbors, here's what you should do..." (They fill in the blanks to be helpful).
  • The French Friend: Says, "That's a tricky situation with many possibilities, but here is a way to think about it..." (They enjoy the ambiguity).
    The robot mimicked these real-world cultural habits perfectly.

4. The "Cultural Blind Spot" (The Big Surprise)

The Finding: This is the most important part. The robot was great at some cultural things but terrible at others.

  • Where it was smart: When asked to write a story in Romanian, it used a specific Romanian folk saying about mushrooms and rain that you wouldn't find in an English book. When asked for help in German, it gave a specific German crisis phone number.
  • Where it was "generic": When asked how to honor a deceased family member in Romanian, it gave a generic, "Western" answer. It missed the specific Romanian religious rituals (like parastas or coliva) that a real Romanian person would know.
    The Analogy: Imagine a chef who can cook a perfect, authentic Italian pizza and a perfect, authentic Japanese sushi. But when you ask for a traditional Romanian stew, the chef just makes a generic "soup" that tastes like it could be from anywhere. The robot has "cultural pockets"—it knows some deep cultural secrets but misses others, depending on the topic.

5. Why This Matters (The "Two-Layer" Test)

The Finding: The author used a two-step method to find these answers.

  • Layer 1 (The Calculator): Counted words and measured how similar the text looked. This told her, "Hey, the Romanian story looks very different from the English one."
  • Layer 2 (The Expert): A human expert (the author) read the answers to ask, "Is this difference a mistake, or is it actually a beautiful cultural choice?"
    The Analogy: A calculator can tell you two paintings have different colors. But only a human art critic can tell you if one painting is "broken" or if it's just a different style of art. The paper argues that we need both the calculator and the human expert to truly understand AI.

The Bottom Line

The paper concludes that AI models like Claude are not just "translators." They are cultural chameleons. They change their personality, length, and depth depending on the language they are speaking. Sometimes this is a beautiful adaptation (like the Romanian story), and sometimes it's a gap (like the missing Romanian funeral traditions).

The author warns that if we only look at "fact-checking" scores, we miss these subtle but important differences. To build fair AI for everyone, we need to listen to how the robot speaks in every language, not just check if it got the facts right.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →