← Latest papers
💬 NLP

Evaluating Modern Large Language Models on Low-Resource and Morphologically Rich Languages:A Cross-Lingual Benchmark Across Cantonese, Japanese, and Turkish

This paper presents a comprehensive cross-lingual benchmark evaluating seven state-of-the-art large language models on Cantonese, Japanese, and Turkish across four diverse tasks, revealing that while top proprietary models lead in performance, significant challenges remain in handling morphological complexity and cultural nuances, particularly for smaller open-source models.

Original authors: Chengxuan Xia, Qianye Wu, Hongbin Guan, Sixuan Tian, Yilun Hao, Xiaoyu Wu

Published 2026-02-13
📖 5 min read🧠 Deep dive

Original authors: Chengxuan Xia, Qianye Wu, Hongbin Guan, Sixuan Tian, Yilun Hao, Xiaoyu Wu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of Artificial Intelligence as a massive, super-smart library. For a long time, this library was filled with books written only in English. The librarians (the AI models) were incredibly well-read in English, able to answer any question, summarize any story, or translate any text with ease.

But what happens when you ask these librarians to help someone who speaks Cantonese, Japanese, or Turkish? Do they just guess? Do they sound like robots? Or have they finally learned to speak these languages with the same fluency and cultural understanding?

This paper is like a report card given to seven of the smartest AI "librarians" in the world. The authors put them through a rigorous test to see how well they handle languages that are either:

  1. Low-Resource: Like Cantonese, where there just aren't many digital books available for the AI to study.
  2. Morphologically Rich: Like Turkish, where words are like Lego bricks that snap together in complex ways to create huge, new words.

The Contestants

The paper tested seven different AI models, ranging from the "Super-Elites" (like GPT-4o and Claude 3.5) to the "Mid-Tier" open-source models (like LLaMA 3.1) and the "Rookies" (smaller models like Mistral 7B).

The Four Challenges (The Exam)

The librarians had to take four different types of tests in each language:

  1. The Trivia Quiz (Open-Domain QA): "Who was the first female pilot in Turkey?" or "What month is this festival in Guangzhou?"
    • The Goal: Can the AI find the right fact without making things up?
  2. The Book Report (Summarization): Reading a long news article and writing a short, clear summary.
    • The Goal: Can the AI keep the main points without losing the flavor of the language?
  3. The Translator (English-to-X): Taking an English sentence and turning it into natural-sounding Turkish, Japanese, or Cantonese.
    • The Goal: Does it sound like a human wrote it, or a machine translating word-for-word?
  4. The Cultural Chat (Dialogue): This was the hardest part. The AI had to chat with a user using local slang, polite honorifics (like in Japanese), or idioms.
    • The Goal: If a user makes a joke in Cantonese, does the AI get the punchline? If a user asks for advice in Turkish, does the AI sound respectful?

The Results: Who Passed?

🏆 The Super-Elites (GPT-4o, GPT-4, Claude 3.5)
These models are like polyglot scholars. They performed the best across the board.

  • GPT-4o was the star of the show, especially with Cantonese. It didn't just translate; it understood the vibe. It knew when to use slang and when to be formal.
  • Claude 3.5 was right behind, showing incredible reasoning skills.
  • The Catch: Even these giants stumbled a little. In Turkish, they sometimes messed up the complex "Lego" word endings. In Cantonese, they occasionally slipped up and used standard Chinese words instead of the local dialect.

🥈 The Mid-Tier (LLaMA 3.1, Mistral Large 2)
These are the hardworking students. They did surprisingly well on Japanese and Turkish, almost catching up to the elites. However, when it came to Cantonese, they struggled. It's like a student who studied hard for a math test but didn't have enough textbooks for the history test. Because there is less data for Cantonese in their training, they often sounded like they were speaking "Standard Chinese" rather than the local dialect.

🥉 The Rookies (LLaMA-2 13B, Mistral 7B)
These models are like fresh graduates who haven't seen much of the world yet. They struggled significantly.

  • They often gave wrong answers to trivia.
  • Their summaries were confusing.
  • In the cultural chat, they were tone-deaf. They might use a casual, rude tone when speaking to a grandmother in Japan, or fail to understand a Turkish proverb entirely.

The Big Takeaways

  1. Size Matters (But Data Matters More): The biggest models generally win, but if a language is "low-resource" (like Cantonese), even the biggest models need more specific training data to truly shine.
  2. Culture is Hard: You can teach an AI grammar, but teaching it culture is like teaching a fish to ride a bicycle. The top models are getting better at understanding local jokes and politeness, but they still make mistakes that a human would never make.
  3. Robots vs. Humans: The paper found that computer programs used to grade these tests (like BLEU scores) often failed. They might give a low score to a perfect translation just because the words were slightly different. Human judges were essential to catch the nuance, humor, and cultural respect that machines miss.

The Bottom Line

We are making amazing progress. The AI librarians are learning to speak more languages than ever before. However, there is still a "digital divide." While English speakers get the best service, speakers of languages like Cantonese, Turkish, and Japanese are still waiting for their AI assistants to truly understand their unique cultures and complex grammar.

The authors released their test data to help everyone build better, more inclusive AI for the whole world, not just the English-speaking half.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →