← Latest papers
💬 NLP

No Optimal Language Set Exists for Multilingual Instruction Tuning: Insights from a Linguistically-Informed Study

This paper systematically evaluates linguistically-informed language selection strategies for multilingual instruction tuning and finds that no universal optimal language set exists, as performance is highly dependent on specific tasks and models, with adding more languages often triggering the curse of multilinguality.

Original authors: Gürkan Soykan, Gözde Gül Şahin

Published 2026-07-28
📖 4 min read☕ Coffee break read

Original authors: Gürkan Soykan, Gözde Gül Şahin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of artificial intelligence as a massive, hungry library. In recent years, scientists have been teaching these digital brains to follow instructions, not just in one language like English, but in dozens or even hundreds. This is called "multilingual instruction tuning." Think of it like training a new employee who needs to speak to customers from every country on Earth. The goal is to make the employee smart enough to understand a request in French, answer in Swahili, and explain a joke in Japanese, all without getting confused.

However, there's a catch. The library has limited shelf space (computing power) and a limited amount of time to train the employee. If you try to cram every single language into the training schedule, the employee might get overwhelmed and start forgetting how to speak any of them well. This is known as the "curse of multilinguality." Because of this, many researchers have wondered: Is there a "magic recipe" for picking the perfect group of languages? Maybe if we pick languages that are geographically spread out, or ones that sound very different from each other, we could create a super-smart employee who works perfectly for everyone, everywhere. It's a hopeful idea: find the perfect mix, and you get the best possible AI.

This paper, titled "No Optimal Language Set Exists for Multilingual Instruction Tuning," goes into the lab to test that exact hope. The researchers, Gürkan Soykan and Gözde Gül Şahin, set up a series of experiments to see if they could find that magic recipe. They took three different types of AI models (think of them as three different trainees with different learning styles) and tried to teach them using various "linguistically-informed" strategies. These strategies were like different ways of picking a team: one team was chosen based on where the languages are spoken on a map (Geography), another based on how the languages are built grammatically (Typology), and others based on how similar their meanings are or how they were learned by computers. They compared these smart picks against a simple "random" pick, like drawing names out of a hat.

The results were a bit of a plot twist. The authors found that there is no single "best" team of languages that wins every time. It's not like finding a universal key that opens every door. Instead, the best choice depended entirely on which specific AI model you were training and what specific task you asked it to do. For one model, picking languages based on geography worked great; for another, picking them based on grammar was better. Sometimes, the "smart" picks actually did worse than just picking languages randomly!

The study also discovered a hard limit. When they kept adding more languages to the training mix, the AI's performance didn't keep getting better forever. Once they hit a certain point—around 13 or 14 languages—the performance actually started to drop. It's like trying to eat a buffet: the first few plates are delicious, but if you keep piling more food on your plate, you eventually get too full to enjoy any of it. This confirmed the "curse of multilinguality": with a fixed amount of training time, adding too many languages hurts the AI's ability to learn them all well.

Furthermore, the researchers found that what works for one AI model might be terrible for another. A strategy that made one model smarter could make a different model dumber. They also noticed that sometimes, teaching the AI in just one language (like Vietnamese) was just as good as, or even better than, teaching it in a mix of many languages, depending on the model.

In the end, the paper suggests that the search for a single, perfect "optimal language set" is a dead end. There is no one-size-fits-all solution. The best approach seems to be a bit more cautious: picking a diverse group of about 10 to 15 languages is usually a safe bet, but you have to be ready to test different combinations for every new model and every new task. The authors warn that simply averaging scores across different tests can be misleading, hiding the fact that a strategy might be great for one job but terrible for another. So, while we can't find a magic recipe, we do know that being thoughtful about how we pick our languages—and knowing when to stop adding more—is the key to building better, more helpful multilingual AI.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →