← Latest papers
💬 NLP

Do Chinese models speak Chinese languages?

This study reveals that despite China's diverse linguistic landscape, its leading open-weight LLMs exhibit multilingual performance profiles strikingly similar to Western models—excelling in major global languages like Mandarin, French, and German while often failing to support minority languages like Kazakh and Uyghur—suggesting that global benchmarking practices and shared training resources are driving a homogenization of language capabilities over localized prioritization.

Original authors: Andrea W Wen-Yi, Unso Eun Seo Jo, David Mimno

Published 2026-05-18
📖 4 min read☕ Coffee break read

Original authors: Andrea W Wen-Yi, Unso Eun Seo Jo, David Mimno

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of Artificial Intelligence as a massive, global cooking competition. Chefs (model developers) from different countries are trying to create the ultimate "Universal Language Pot" that can understand and speak every language on Earth.

This paper asks a specific question: Do the chefs from China, who have a kitchen full of diverse local ingredients, actually cook with them? Or do they just follow the same recipe as the chefs from the US and Europe?

Here is the breakdown of their findings, using simple analogies:

1. The Setup: A Diverse Kitchen vs. A Global Menu

China is like a massive kitchen with 1.4 billion people speaking Mandarin, plus hundreds of other local dialects and minority languages (like Uyghur, Tibetan, and Kazakh). You would expect the Chinese chefs to have a special advantage: they have easy access to these local ingredients.

However, there is a catch. To win the competition and get famous globally, the chefs are judged by a panel of critics who mostly speak English and use English-based scorecards (benchmarks). If you want to be recognized as a top chef, you have to make your dish taste great to these English-speaking judges.

2. The Experiment: Tasting the Dishes

The researchers tasted dishes from 6 Chinese chefs (like Qwen, Yi, and DeepSeek) and 4 Western chefs (like Llama and Mistral). They tested them on 21 different "languages," ranging from major ones like French and German to local Chinese minority languages.

They used three different ways to taste the food:

  • The Translation Test: Can the model translate a sentence from English to another language without losing the flavor (meaning)?
  • The Reading Test: Can the model read a story in a foreign language and answer questions about it correctly?
  • The Name Game: Can the model look at a text and correctly say, "This is written in Kazakh" or "This is written in Tibetan"?

3. The Results: The "Mandarin Exception"

The results were surprisingly uniform, with one big exception.

  • The "Clone" Effect: For almost every language tested (French, German, Japanese, Korean, etc.), the Chinese models and the Western models performed almost identically. It's as if both groups of chefs were following the exact same global recipe book. The paper found a 93% correlation between them. If a Western model was good at French, the Chinese model was also good at French. If a Western model struggled with a minority language, the Chinese model struggled just as much.
  • The Mandarin Superpower: The only time the Chinese chefs truly outshined the Western ones was with Mandarin Chinese. Because they had more access to Mandarin data and focused on it, their "Mandarin dish" was significantly tastier than the Western chefs' version.
  • The Minority Language Blind Spot: Here is the sad part. Despite having access to local ingredients, the Chinese models were just as bad at handling Chinese minority languages (like Uyghur and Kazakh) as the Western models were. In fact, many Chinese models couldn't even recognize that a text was written in Uyghur or Kazakh. They treated these local languages as if they didn't exist, just like the Western models did.

4. Why Did This Happen?

The paper suggests two main reasons:

  1. The "Global Scorecard" Pressure: To get famous and get funding, Chinese developers are obsessed with winning global benchmarks. These benchmarks are mostly in English or major European languages. So, instead of using their unique access to local Chinese data to build a truly diverse model, they optimized their models to look good on the global scoreboard.
  2. The "Mandarin-First" Policy: While China has a history of trying to protect its many languages, recent policies have focused heavily on making Mandarin the single, unifying language for the whole country. The AI developers seem to be following this trend: they prioritize Mandarin (the "super language") and ignore the rest.

The Bottom Line

The paper concludes that despite China having a unique linguistic landscape and the potential to build a model that truly speaks to its diverse population, the Chinese models have become "homogenized."

They are essentially carbon copies of Western models, with one small upgrade: they speak Mandarin slightly better. They have not used their unique position to save or elevate the hundreds of other languages spoken within China's borders. Instead, they are all cooking the same global dish, leaving the local, minority flavors on the shelf.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →