← Latest papers
💻 computer science

Are Multilingual Models Actually Improving? Isolating True Cross-Lingual Transfer

This paper introduces the Hardness Adjusted Transfer (HAT) Score to isolate true cross-lingual transfer from general source-language improvements, revealing that while small models transfer effectively and progress has been made over time, scaling models yields slower-than-expected gains in transfer strength.

Original authors: Prasoon Bajpai, Eleftheria Briakou, Colin Cherry, Preethi Jyothi, Vihari Piratla

Published 2026-06-23
📖 5 min read🧠 Deep dive

Original authors: Prasoon Bajpai, Eleftheria Briakou, Colin Cherry, Preethi Jyothi, Vihari Piratla

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Faking It" Trap

Imagine you are a teacher grading students from two different countries: Country A (where the teacher speaks the language fluently) and Country B (where the teacher is still learning).

For a long time, researchers measured how well a computer model learned a new language by simply looking at its test scores in that new language. They thought: "If the score goes up, the model is getting better at transferring its knowledge!"

The paper argues this is a trick.

The authors say that often, the model isn't actually getting better at translating its knowledge; it's just getting smarter overall.

  • The Analogy: Imagine a student who is terrible at math but gets a tutor. The tutor helps them understand the concepts better. Now, when the student takes a test in their native language, they get a 90%. When they take the same test in a foreign language, they get a 60%.
  • Later, the student gets a better tutor. They understand the concepts even deeper. Now they get a 98% in their native language and a 65% in the foreign language.
  • The Mistake: If you just look at the foreign language score (60% → 65%), you might think the student improved their language skills. But really, they just got better at the math. The "gap" between their native and foreign scores actually got worse (from 30 points down to 33 points).

The paper says current methods are like that: they confuse "getting smarter generally" with "getting better at cross-language transfer."

The Solution: The "HAT Score" (Hardness Adjusted Transfer)

To fix this, the authors invented a new way to measure progress called the HAT Score.

The Analogy: Imagine a "Perfect Transfer" line. This is a magic line where if a student gets 80% in their native language, they should get 80% in the foreign language if they have truly mastered the transfer.

  • Old Method: Just looked at the final number (e.g., "65%").
  • HAT Score: Looks at the relationship. It asks: "Given that the model got X% in the source language, how much higher (or lower) did it perform in the target language compared to what we expected?"

If a model performs above the expected line, it's a true transfer win. If it just follows the line because it got smarter generally, the HAT score stays the same. It filters out the "noise" of general improvement to find the "signal" of true language learning.

What They Found (The Results)

The researchers tested 20 different AI models (from companies like Google, OpenAI, and Anthropic) on three different types of tasks. Here is what their "HAT Score" revealed:

1. Small Models Aren't Broken

  • The Myth: People thought small AI models were terrible at learning new languages.
  • The Reality: When you use the HAT Score, small models actually do a decent job! They aren't "broken"; they just have lower overall scores. The old metrics made them look worse than they really were.

2. Progress is Slower Than We Thought

  • The Myth: As AI models get bigger (more parameters), they magically become perfect at cross-language transfer.
  • The Reality: Using the HAT Score, the authors found that making models bigger helps, but the improvement is much slower than we hoped. We aren't seeing the massive leaps in language transfer that the raw scores suggested.

3. Time is Helping (But Only on Some Tasks)

  • The Reality: Newer models (released in 2025/2026) are genuinely better than older ones at reasoning tasks (like math and logic puzzles). On these tests, the HAT Score shows they are almost "solved"—meaning they transfer knowledge almost perfectly.
  • The Catch: This progress hasn't happened for fact-recall tasks (like answering trivia questions). On those, progress has stalled.

4. The "Script" Barrier

  • The Reality: Does it matter if the languages use different alphabets? (e.g., English vs. Arabic).
  • The Finding: It depends on the task.
    • For reasoning (math/logic), the alphabet doesn't matter much. The model figures it out.
    • For facts (trivia), switching alphabets creates a huge barrier. The model struggles much more when the script changes.

5. "Thinking" Helps

  • The Reality: Models that are allowed to "think" (generate a chain of reasoning before answering) perform better at cross-language transfer.
  • The Observation: The models seem to use this extra thinking time to bridge the gap between languages, effectively translating the problem in their "mind" before solving it.

The Conclusion

The paper concludes that while AI is getting better, we have been celebrating "general smarts" as if they were "language skills."

  • The Good News: We are making real progress on reasoning tasks, and small models are surprisingly capable.
  • The Bad News: We are making slower progress on fact-based tasks, and the gap between languages is still stubborn, especially when different alphabets are involved.
  • The Call to Action: The current tests (benchmarks) are too easy for the newest models. We need harder, more complex tests that force models to struggle with multiple steps, where the "translation" part of the problem is the real challenge.

In short: Don't just look at the final score; look at how the model got there. The HAT Score is the new ruler for measuring true language learning.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →