← Latest papers
💬 NLP

Skill Issue: Are Skills Language-Invariant in LLMs?

This paper demonstrates that large language models exhibit significant, measurable skill inconsistencies across different languages by using a multilingual self-play framework to show that the same model can display markedly different performance strengths, reasoning capabilities, and strategic behaviors depending solely on the language interface used.

Original authors: Bobby Cheng, Adam Gaber, Zhengyuan Liu, Catherine Arnett, Omer Goldman, Cheston Tan, Leshem Choshen

Published 2026-08-27
📖 5 min read🧠 Deep dive

Original authors: Bobby Cheng, Adam Gaber, Zhengyuan Liu, Catherine Arnett, Omer Goldman, Cheston Tan, Leshem Choshen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Large language models are the engines behind many of the artificial intelligence tools we use today, capable of writing stories, solving problems, and answering questions in dozens of languages. For years, researchers have known that these systems do not perform equally well in every tongue; a model might be brilliant at answering a history question in English but struggle to find the same fact when asked in Arabic or Chinese. This inconsistency has usually been blamed on the data the models were trained on, assuming that some languages simply have less information available to teach the machine. However, a new line of inquiry asks a deeper question: is the problem just about what the model knows, or is it about what the model can actually do? Could it be that the same artificial mind possesses different sets of skills depending on which language it is speaking, even when the task itself remains exactly the same?

A team of researchers set out to answer this by treating language not as a subject of study, but as a variable in a controlled experiment. They turned to a familiar concept: the game. By placing two identical copies of the same artificial intelligence model into a digital arena to play against each other, they created a perfect test of skill. In this setup, the rules of the game, the board, the cards, and the possible moves remained frozen and identical for both players. The only thing that changed was the language used to describe the situation. One player received instructions and saw the game state in English, while its opponent saw the exact same board and rules translated into German, Hebrew, Chinese, or one of several other languages. Because the models were identical and the game mechanics were fixed, any difference in who won had to come from the language interface itself, not from a difference in knowledge or intelligence.

The researchers built a massive testing ground called a multilingual extension of a game platform known as TextArena, which allowed them to run thousands of matches across six different types of games. These games ranged from simple spatial puzzles like Tic-Tac-Toe, where players must visualize a grid, to complex strategic games involving resource allocation and bluffing. They tested three different open-source models, each with a similar size, pitting them against themselves in eight different languages. The results revealed a startling reality: the same model could be a master strategist in one language and a clumsy novice in another. In some cases, the model playing in English consistently defeated its own version playing in Hebrew or Arabic, not because it knew more about the game, but because the language itself seemed to unlock or block its ability to think clearly.

The nature of these failures was specific and revealing. In games that required spatial reasoning, such as connecting lines on a grid, models playing in languages that use non-Latin scripts, like Arabic or Hebrew, frequently made errors that looked like a fundamental misunderstanding of the board. They would often fail to see a winning line that ran diagonally or vertically, treating the grid as if the rows were the only thing that mattered. In games of chance and strategy, like a simplified form of poker, the models changed their entire personality based on the language. A model might play cautiously and honestly in one language but become aggressive and prone to risky bluffs in another, even though the cards in its hand were identical. The researchers found that these were not random mistakes; they were systematic shifts in behavior that depended entirely on the words used to frame the problem.

Perhaps the most surprising discovery was that the problem was not always about the language the model was reading, but the language it used to think. In some experiments, the researchers kept the game instructions in a weaker language, like German, but instructed the model to perform its internal reasoning in English. This simple switch allowed the model to recover much of its lost performance, suggesting that the language barrier was not just a translation issue but a cognitive one. The model could access its full strategic potential if it was allowed to "think" in a stronger language, even while playing in a weaker one. This indicates that the language interface affects the very stages of decision-making, from understanding the current state of the game to selecting the best move.

The study also looked at whether these differences could be explained by how much text exists on the internet for each language. While it is true that languages with more available data generally performed better, this factor did not tell the whole story. Some languages with relatively little data on the web still allowed the models to play very well, while others with abundant data resulted in poor performance. This suggests that the issue is not just about the quantity of training material, but about how the model has learned to connect skills across different linguistic systems. The findings imply that for artificial intelligence to be truly reliable and fair, developers cannot simply assume that a model's abilities are consistent across all languages. The language a model speaks is not just a window into its knowledge; it is a lens that can distort its skills, changing how it sees the world and how it plays the game.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →