A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures
This paper systematically evaluates four multilingual interpretability methods across 21 models, finding that anisotropy distorts most metrics while ILO uniquely and robustly correlates with cross-lingual transfer performance, leading the authors to recommend ILO as the primary sharing metric alongside anisotropy diagnostics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Modern artificial intelligence systems that understand language are built on a foundation of mathematics that maps words and sentences into vast, multi-dimensional spaces. In these spaces, the meaning of a word is represented by a specific point, and the relationship between words is defined by the distance and direction between those points. When these models are trained on many different languages at once, they develop a remarkable ability: they begin to use the same internal points to represent similar concepts across those languages, effectively creating a shared mental dictionary. This phenomenon, known as cross-lingual sharing, is the engine that allows a model trained mostly on English to answer questions in Swahili or translate poetry from Japanese. However, for years, scientists have struggled to measure exactly how well this sharing happens. Different research teams have built different rulers to measure the same thing, and often, those rulers gave conflicting answers. One team might say a model shares its knowledge perfectly, while another, using a different method, might say it barely shares anything at all. This confusion made it difficult to know which models were truly multilingual and which were just pretending.
A team of researchers set out to resolve this confusion by comparing four representative measurement tools selected to contrast different methodological families against a wide variety of language models. They gathered twenty-one different models, ranging from small, efficient versions to massive, complex ones, representing five distinct families of artificial intelligence. These models varied in size from 125 million to 14 billion parameters, a scale that captures the diversity of the field. The researchers fed these models a standard set of sentences in ten different languages, including Arabic, Chinese, English, and Turkish, and then applied each of the four measurement tools to see how much the models mixed their languages together. They also tested the models' real-world performance by teaching them a new task in English and seeing how well they could apply that knowledge to other languages without further training. This functional test served as a validation criterion, a way to see if the measurements actually predicted the models' ability to transfer knowledge, though the authors note that transfer is driven by factors beyond just representational alignment.
The results revealed that the disagreement between the measuring tools was not a reflection of the models themselves, but rather a flaw in how the tools were built. The researchers discovered that the internal representations of these models are often "anisotropic," a term describing a shape that is stretched into a narrow, needle-like cone rather than spreading out evenly in all directions. Imagine the difference between a sphere and a thin pencil; in the pencil-shaped space, almost any two points look close to each other simply because the space is so narrow. Three of the four measuring tools were fooled by this shape. They interpreted the fact that all points were crowded together as evidence of strong sharing, even when the models were not actually mixing their languages. One tool, in particular, relied on a method that stripped away the most important directions of information to find patterns, only to discover that the language identity had been pushed into the discarded directions. It was like trying to find a specific flavor in a soup by removing the salt, only to conclude the soup was bland because the salt was gone.
In contrast, one specific measurement, called Interlingual Local Overlap, proved to be the only reliable ruler. This tool looked at the immediate neighbors of each word in the model's internal space, asking whether words from different languages were standing next to each other. Unlike the other tools, this method survived rigorous testing. It consistently predicted which models would perform well on cross-lingual tasks, showing a strong correlation that held true regardless of the model's size or family. The researchers found that models with high scores on this specific metric were the ones that could successfully transfer knowledge from English to other languages. The other tools either failed to distinguish between models that shared knowledge and those that did not, or they were so influenced by the narrow shape of the data that they produced misleadingly high scores for models that were actually quite poor at sharing.
The study concludes that the field has been misled by tools that are sensitive to the geometric distortions of the data rather than the actual mixing of languages. The researchers recommend that future studies use Interlingual Local Overlap as the primary sharing metric, but to be reported alongside anisotropy diagnostics. Furthermore, they suggest that whenever anyone reports a score for language sharing, they must also report a diagnostic check for the shape of the data, ensuring that a high score reflects genuine mixing and not just a narrow, distorted space. This work does not claim to have solved the mystery of how language models learn, but it provides a clear, trustworthy way to measure what they have learned. By identifying the tools that work and the ones that fail, the researchers have cleared the fog, allowing the scientific community to finally compare models on a level playing field and understand the true nature of multilingual intelligence. These results are corrective rather than dismissive of global sharing metrics: a single global score remains the right instrument for quantifying and comparing sharing across models, while explaining how sharing is implemented calls for mechanistic methods.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.