The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty
This paper quantifies a "tokenizer tax" across 25 European languages, revealing that non-English languages, particularly Ukrainian and those with complex morphology, suffer from significantly higher token-to-word ratios due to underrepresentation in pre-training data, while demonstrating that these fertility rankings remain consistent across domains and that few-shot effects are model-intrinsic rather than language-dependent.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are paying for a taxi ride. In English, the taxi charges you by the "stop" (a word). But in Ukrainian, the taxi driver insists on charging you by the "step" (a tiny piece of a word). Because Ukrainian words are often longer and more complex, the driver takes many more steps to get to the same destination. You end up paying twice as much for the exact same ride, even though the distance hasn't changed.
This paper, written by Volodymyr Ovcharov, is all about that hidden "tax" on non-English languages when using AI. Here is the breakdown of what they found, using simple analogies:
1. The "Step" Tax (Tokenizer Fertility)
AI models don't read whole words like humans do. They chop text into tiny chunks called tokens.
- The Analogy: Think of a word as a loaf of bread.
- English: The AI cuts the loaf into big, thick slices. You only need 1.2 slices to represent one word.
- Ukrainian: The AI uses a dull knife and cuts the same loaf into tiny, crumbly pieces. It takes 2.7 slices to represent one word.
- The Cost: Since AI companies charge you per "slice" (token), speaking Ukrainian costs you roughly 2.2 times more than speaking English for the same amount of text.
- The Hierarchy: The paper measured 25 European languages.
- The Cheap Group: English, Spanish, and French (1.2 to 1.7 slices per word).
- The Middle Group: German and Dutch (1.7 to 1.9 slices).
- The Expensive Group: Slavic languages like Polish and Czech (2.2 to 2.5 slices).
- The Most Expensive: Ukrainian, Greek, and Maltese (around 3.1 slices).
2. The "Ukrainian Penalty"
The researchers found something surprising about Ukrainian. Even though Ukrainian, Polish, and Czech are all Slavic cousins with similar word structures, Ukrainian is significantly more expensive to process.
- The Analogy: Imagine two identical factories. One (Polish) has a well-stocked supply of pre-cut bricks because it's been building for a long time. The other (Ukrainian) has to cut every brick from raw stone because it was ignored in the past.
- The Cause: The paper shows that Ukrainian has 2 to 6 times less training data in the AI's "library" compared to Polish or Czech. Because the AI hasn't seen Ukrainian words as often, it doesn't know how to group them efficiently. It keeps chopping them up into tiny, inefficient pieces.
3. The "One-Size-Fits-All" Knife
The paper tested if this "tax" changes depending on what you are talking about (e.g., legal contracts vs. news vs. Wikipedia).
- The Finding: It doesn't matter. The ranking of which AI is "cheap" and which is "expensive" stays exactly the same whether you are talking about football, law, or history.
- The Takeaway: If you test an AI on one type of text, you know exactly how much it will cost you for any other type of text. You don't need to re-test it every time.
4. The "Magic Trick" (Few-Shot Learning)
AI can sometimes learn a new task just by seeing a few examples (like showing it a sample question and answer). The researchers tested if this "magic trick" worked differently for different languages.
- The Finding: The trick isn't about the language; it's about the AI's personality (its design).
- Some AIs (like "Nemotron") get smarter when shown examples, no matter if the text is Ukrainian, Polish, or Russian.
- Other AIs (like "Maverick") actually get dumber when shown examples, again, regardless of the language.
- The Lesson: Don't blame the language for the AI's confusion. Blame the AI's design. If an AI fails with examples in English, it will likely fail with examples in Ukrainian too.
5. Why Some AIs Are Better at Chopping
The researchers looked under the hood to see how different AIs cut the words.
- The Good Choppers: Some AIs (like Gemma 2) cut Ukrainian words at natural "morpheme" boundaries (the meaningful parts of a word, like roots and endings). This is efficient.
- The Bad Choppers: Other AIs (like Qwen3) cut words at random spots, sometimes splitting a single meaningful part in half. It's like cutting a sandwich right through the middle of a pickle instead of between the slices of bread.
- The Surprise: Having a bigger dictionary (vocabulary) doesn't guarantee better cutting. Some AIs with huge dictionaries still chopped Ukrainian words into tiny, inefficient pieces because they just didn't have enough Ukrainian examples in their training data to learn the right cuts.
Summary of Recommendations
The paper suggests three simple rules for anyone using AI with Ukrainian or similar languages:
- Expect the Tax: Budget for the fact that Ukrainian will cost about double what English costs.
- Test Once: You only need to measure the "tax" once; it applies to all topics.
- Be Careful with Examples: Don't assume that showing an AI examples will help. For some models, it might actually hurt performance, and this happens in every language.
The Bottom Line: The AI world is currently built with an English bias. Until the "libraries" of training data are balanced, languages like Ukrainian will pay a hidden "step tax" just to be understood.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.