Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures
This paper demonstrates that Large Language Models generalize better across languages when their latent representations accurately capture the hierarchical similarity structures of language families, such as the Indo-European tree, with this structural alignment directly correlating to improved performance on multilingual benchmarks like XNLI.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to understand the world. You don't just give it a dictionary; you let it read millions of books, websites, and stories. Over time, the robot starts to build a mental map of how things relate to one another. In the world of computer science, this is called a "Large Language Model" (LLM). Think of these models as super-smart digital brains that have read almost everything on the internet. But here's the tricky part: most of what they read is in English. When you ask them to speak a language they haven't seen much of, like Swahili or Icelandic, they often stumble.
Scientists have long wondered why this happens. A key idea from psychology suggests that our brains (and smart computers) learn by noticing similarities. If you know what a "dog" is, you can guess what a "wolf" is because they look and act alike. This is called "generalization." The better your mental map groups similar things together, the better you are at guessing new things. But does this same rule apply to entire languages? Do these digital brains secretly organize languages the way a family tree organizes relatives? That is the big question this paper sets out to answer.
The Great Language Family Reunion
Picture a giant, invisible party where every language in the Indo-European family (a huge group including English, Spanish, Hindi, and Russian) is invited. In the real world, we know these languages are related like cousins, siblings, and distant relatives. Some, like Spanish and Italian, are like twins who grew up in the same house. Others, like English and Hindi, are distant cousins who haven't seen each other in thousands of years.
The researchers in this paper wanted to see if Large Language Models (LLMs) secretly know this family tree. They didn't ask the models, "Hey, are these languages related?" Instead, they looked at the models' "thoughts" (their internal math) to see how they grouped languages together. It's like watching a guest at the party and seeing who they stand next to. If the model puts Spanish and Italian close together, but keeps them far away from Hindi, it's showing it understands the family connections without ever being explicitly taught them.
The Discovery: The Models Have a Secret Map
The team tested 12 different AI models, ranging from older ones to newer, more powerful ones. They fed the models thousands of sentences in 38 different languages. Then, they used a special mathematical trick to map out how the models "felt" about each language.
The result was surprising and delightful: The models had accidentally built a map that looked almost exactly like the real family tree of languages.
Just like a genealogist drawing a family tree, the models naturally grouped languages into the same sub-families: Germanic languages (like English and German) huddled together; Romance languages (like French and Spanish) formed their own circle; and Slavic languages (like Russian and Polish) stuck to their own group. The models didn't need a teacher to tell them, "These are Romance languages." They figured it out on their own just by reading the data.
Why Some Models Are Better at the Party
But here is the twist: not all models are equally good at this. The researchers found a direct link between how well a model organized its language map and how well it actually performed on a difficult test called XNLI (a test where the model has to guess if one sentence proves, contradicts, or has nothing to do with another sentence in a different language).
Think of it like this: If a student has a messy, jumbled notebook where they mix up all their history facts, they will struggle to answer a test question. But if their notebook is perfectly organized by topic, they can find the answer instantly. The paper found that models which organized their "language notebook" correctly (keeping similar languages close together) were much better at solving the test problems.
Specifically, the models that were trained on data from many different languages (multilingual models) did a much better job of drawing this family tree than models trained mostly on English. It's as if the multilingual models had a bigger, clearer map, while the English-only models were trying to navigate with a blurry, incomplete sketch.
The "Script" Trap and Other Surprises
One might guess that the models just grouped languages by how they look on paper (their alphabet or script). For example, maybe they thought Hindi and Urdu were similar just because they both use the Devanagari script? The paper says no. While the way words are written (orthography) does have a small effect, it's not the main reason. The models were smart enough to see that Hindi and Urdu are linguistic cousins, even though they use different scripts, and that Persian (which uses an Arabic script) is actually related to North Indian languages, not Arabic ones. The models were looking at the structure of the language, not just the font.
However, the models weren't perfect. They sometimes got a bit confused with "orphan" languages like Greek, Welsh, and Albanian. In some models, Greek and Welsh were paired up strangely, even though they aren't closely related. This suggests that while the models are great at seeing the big picture, they sometimes trip over the unique, isolated branches of the family tree.
When Did They Learn This?
The researchers also watched the models as they were being trained, like watching a baby learn to walk. They found that the models started organizing languages into these family groups very early in their training—after just 100 steps of learning. This suggests that figuring out "who is related to whom" is one of the first things a smart language model learns, before it even gets into the nitty-gritty details of specific words.
The Bottom Line
This paper suggests that the secret to a language model's success isn't just memorizing more words; it's about building a good mental map of how languages relate to each other. When a model understands that Spanish and Italian are close cousins, and that English and German are siblings, it becomes much better at translating ideas from one to the other.
It turns out that these digital brains are surprisingly good at history. They didn't need a textbook to learn the Indo-European family tree; they just needed to read the world, and the patterns emerged on their own. And the better they get at drawing that map, the better they get at speaking every language on it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.