← Latest papers
💬 NLP

Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages

This paper introduces a human-validated framework revealing that large language models exhibit significantly higher reasoning-conclusion misalignment in non-Latin scripts compared to Latin ones, despite high task accuracy, thereby exposing critical blind spots in current multilingual evaluation practices.

Original authors: Anaelia Ovalle, Candace Ross, Sebastian Ruder, Adina Williams, Karen Ullrich, Mark Ibrahim, Levent Sagun

Published 2026-03-31
📖 4 min read☕ Coffee break read

Original authors: Anaelia Ovalle, Candace Ross, Sebastian Ruder, Adina Williams, Karen Ullrich, Mark Ibrahim, Levent Sagun

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a student to take a history test. You give them the same exam in English, Spanish, Hindi, and Korean.

The Old Way of Checking:
In the past, if the student got the right answer on the Korean section, you would say, "Great job! They understand history." You only looked at the final bubble they filled in on the answer sheet.

The New Discovery:
This paper says: "Wait a minute. Let's look at how they got there."

The researchers built a system to read the student's "scratch work" (their reasoning) and check if the logic actually leads to the answer. They found a shocking blind spot: The student is often getting the right answers for the wrong reasons, especially when the test is in a language they aren't as fluent in.

Here is the breakdown using simple analogies:

1. The "Lucky Guess" vs. The "Real Logic"

Imagine a student taking a math test.

  • Scenario A (English): They write out the steps: "2 + 2 = 4. 4 + 4 = 8." Then they circle 8.
    • Verdict: Perfect. The logic supports the answer.
  • Scenario B (Korean): They write: "I think the answer is 8 because... uh... numbers are cool? Also, 2+2 is 4." Then they circle 8.
    • Verdict: The answer is right, but the logic is nonsense. They got lucky.

The paper found that for languages using non-Latin scripts (like Korean, Arabic, Hindi), the models were much more likely to be in "Scenario B." They were guessing the right answer but couldn't explain why in a way that made sense.

2. The "Script" Problem

The researchers noticed something weird about the shape of the letters.

  • Latin Scripts (English, Spanish): The models were like a reliable GPS. They gave the right turn-by-turn directions to get to the destination.
  • Non-Latin Scripts (Korean, Arabic, Hindi): The models were like a drunk friend giving directions. They might point you to the right building (the correct answer), but the path they described was full of holes, contradictions, or made-up facts.

In fact, the "drunk friend" behavior was twice as common in non-Latin languages compared to Latin ones. It wasn't just that the models knew less; it was that their thinking process broke down differently.

3. The "Magic 8-Ball" Effect

The paper calls this Reasoning-Answer Misalignment.
Think of the AI as a Magic 8-Ball.

  • If you shake it in English, it says "Outlook good" and gives you a logical reason why.
  • If you shake it in Korean, it might still say "Outlook good" (the correct answer), but if you look inside the ball, the reason it gives is "Because the stars aligned" or "Because I feel like it."

The AI is essentially hallucinating a justification after it has already guessed the answer. It's not thinking to the answer; it's guessing the answer and then inventing a story to make it look smart.

4. Why Does This Matter?

You might ask, "If the answer is right, who cares how they got there?"

The paper argues that trust is the issue.

  • If a doctor uses an AI to diagnose a disease, and the AI says "You have the flu" (Correct), but the reasoning is "Because you look tired and I like the flu," that's dangerous.
  • If the AI is used for legal or political decisions in different languages, and it's "guessing" in those languages while "thinking" in English, we are building a system that is unfair and unreliable for half the world.

The Takeaway

The researchers created a new "report card" for AI. Instead of just grading the final answer (A, B, C, or D), they grade the essay the AI writes to get there.

Their main finding: Current AI models are great at English reasoning, but when they switch to other languages, they often stop "thinking" and start "pattern matching." They get the right answer by luck or memorization, but their internal logic falls apart.

In short: Just because the AI gives you the right answer doesn't mean it actually understands the question. And right now, it understands English much better than it understands the rest of the world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →