Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap
This paper demonstrates that fixed output-token caps act as a hidden experimental variable that artificially inflates or reverses measured multilingual reasoning gaps, arguing that evaluations must report accuracy across a spectrum of output budgets rather than at a single cap to distinguish true reasoning deficits from truncation artifacts.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to measure how fast two different runners can finish a race. One runner speaks a language where words are short and punchy, while the other speaks a language where words are long and flowery. If you tell both runners, "You must stop running exactly when you have taken 100 steps," you aren't just testing their speed; you are secretly testing how many steps their language requires to say the same thing. The runner with the long words might get cut off mid-sentence, not because they are slow, but because the finish line was set too early for their specific stride. This is the core puzzle of modern Artificial Intelligence: when we ask AI models to solve math problems in different languages, we often see them perform worse in their native tongue than when we ask them to translate the problem into English first. Scientists call this the "multilingual reasoning gap." But is this gap a sign that the AI is actually "dumber" in its native language, or is it just a measurement error caused by how we set the rules of the game?
This paper, titled "Mind the Cap," investigates whether that famous gap is actually a trick of the measurement tape. The researchers set up a race for two popular AI models (Qwen and Llama) to solve math problems in German, Thai, and Swahili. They compared two strategies: letting the AI think and answer in the native language (NATIVE) versus translating the problem to English, thinking in English, and then answering in the native language (TRANSLATE-ACT). The twist? They didn't just look at the final score; they watched the race happen step-by-step, changing the "token cap"—the maximum number of digital building blocks (tokens) the AI is allowed to use to write its answer.
The authors discovered that the "gap" isn't a fixed truth; it's a moving target that depends entirely on how tight the leash is. When the cap is very tight (like 128 tokens), the native language often loses badly, not because the AI can't reason, but because it simply runs out of space to finish its sentence. However, as they loosened the cap, the gap shrank and sometimes even vanished. In fact, at a specific budget of 1,024 tokens, the gap for one model (Qwen) disappeared almost entirely, suggesting that the AI wasn't lacking reasoning skills at all; it just needed more room to breathe.
The study also found something surprising: telling the AI what the limit is before it starts can change how it behaves. When they told the Thai-native AI, "You only have 128 tokens," it actually performed better than when they told it, "You have 2,048 tokens," even though the hard limit was the same in both cases. It seems the AI tries to be more efficient when it thinks it's in a hurry.
Ultimately, the paper argues that the "reasoning gap" is often just a "budget artifact." If you give the AI enough space to finish its thoughts, the native language often catches up. The researchers tested this by freezing their experiments in advance and running thousands of new, independent tests. They found that while the gap is real at tight budgets, it often disappears once the AI is allowed to write a full answer. They also tested a "vocabulary extension"—giving the AI new, shorter words for its native language—but found this only helped when the budget was still tight enough to cut off the AI's sentences. Once the AI had enough space, the extra words didn't make a difference.
The big takeaway is that we shouldn't treat the token limit as a background setting. It's a variable that changes the outcome. If we want to know if an AI can truly reason in a language, we can't just look at a single score with a single limit. We have to watch the whole race, from the starting gun to the finish line, and realize that sometimes, the AI isn't failing to think; it's just failing to finish.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.