← Latest papers
💬 NLP

SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic Languages

This paper introduces SiniticMTError, a novel fine-grained dataset featuring error span, type, and severity annotations for machine-translated examples across several Sinitic languages (Mandarin, Cantonese, Wu, and Hokkien), designed to advance translation quality estimation and error-aware generation while highlighting the current limitations of large language models in detecting such errors.

Original authors: Hannah Liu, Junghyun Min, En-Shiun Annie Lee, Ethan Yue Heng Cheung, Shou-Yi Hung, Elsie Chan, Shiyao Qian, Runtong Liang, Kimlan Huynh, Wing Yu Yip, York Hay Ng, TSZ Fung Yau, Ka Ieng Charlotte Lo, Y
Published 2026-03-18
📖 5 min read🧠 Deep dive

Original authors: Hannah Liu, Junghyun Min, En-Shiun Annie Lee, Ethan Yue Heng Cheung, Shou-Yi Hung, Elsie Chan, Shiyao Qian, Runtong Liang, Kimlan Huynh, Wing Yu Yip, York Hay Ng, TSZ Fung Yau, Ka Ieng Charlotte Lo, You-Wei Wu, Richard Tzong-Han Tsai

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, multilingual robot translator. It's great at translating between big, popular languages like English and Mandarin. But when you ask it to translate into smaller, regional languages like Cantonese, Wu (Shanghai dialect), or Hokkien, it starts to stumble. It might get the general idea right, but it messes up the details, sounding awkward or even making up words that don't belong.

This paper introduces a new tool called SINITICMTERROR. Think of it as a "Mistake Map" or a "Training Manual for Errors" designed specifically to teach AI how to spot and fix its own translation mistakes in these under-served Chinese languages.

Here is a breakdown of what the researchers did, using some everyday analogies:

1. The Problem: The "Overconfident Tourist"

Imagine a tourist who knows a little bit of a local language. They can order food and ask for directions, but if you ask them to describe the weather, they might say, "The sky is very beautiful" (which is grammatically okay but sounds weird in that culture, where you'd just say "The weather is nice").

Current AI models are like that tourist. They are fluent enough to sound natural, but they often make subtle, "silly" mistakes:

  • Adding things that aren't there: Like the tourist adding "I think" when the original sentence didn't include it.
  • Using the wrong word: Translating "beautiful" for a weather report when the local language expects "good."
  • Missing the point: Leaving out crucial words entirely.

For languages like Cantonese, Wu, and Hokkien, these mistakes are common because there isn't enough high-quality data to train the AI properly.

2. The Solution: The "Mistake Map" (SINITICMTERROR)

The researchers built a massive dataset where they took thousands of sentences, had an AI translate them, and then hired native speakers (the "local experts") to go through the translations with a red pen.

They didn't just say "this is wrong." They created a detailed log for every single error:

  • Where is the error? (Pinpointing the exact words).
  • What kind of error is it? (Did the AI add a word? Did it forget a word? Did it use the wrong grammar?).
  • How bad is it? (Is it a minor typo, or did it completely change the meaning?).

The Analogy: Imagine a teacher grading a student's essay. Instead of just giving a "C," the teacher highlights every specific sentence, writes "Grammar error here," "Too formal here," and "You missed a key detail there." SINITICMTERROR is that detailed grading sheet, but for AI.

3. The Languages: The "Family Reunion"

The paper focuses on four specific branches of the "Sinitic" (Chinese) language family:

  • Mandarin: The big, popular cousin that everyone knows.
  • Cantonese: The rich, expressive cousin spoken in Guangdong and Hong Kong.
  • Wu (Shanghainese): The sophisticated cousin from Shanghai.
  • Hokkien: The ancient, dialect-heavy cousin spoken in Fujian, Taiwan, and Southeast Asia.

The researchers found that while Mandarin is getting better, the other three are still struggling. The AI often treats them like "broken" versions of Mandarin, rather than unique languages with their own rules. For example, Cantonese has many tiny "particles" at the end of sentences to show attitude, and the AI keeps dropping them or using the wrong ones.

4. The Test: "Can the AI Grade Itself?"

The researchers didn't just make the map; they tested if current super-smart AI models (like GPT-4 or Gemini) could use this map to find errors.

The Result: It was a bit of a disaster. Even the smartest AI models only got about 30% to 40% of the errors right.

  • The Metaphor: It's like asking a brilliant math student to grade a test on a subject they've never studied. They can see something is wrong, but they can't pinpoint exactly what or how bad it is. They often miss the subtle cultural nuances that a native speaker would catch immediately.

5. Why This Matters

This dataset is a game-changer for two reasons:

  1. It's a Training Gym: Now, developers can use this "Mistake Map" to train their AI models to become better at spotting errors, much like a coach using game tape to show a player where they missed the ball.
  2. It Gives a Voice to the "Small" Languages: By treating these languages with the same rigorous attention as English or Mandarin, the paper argues that these languages deserve better technology. It highlights that current AI is biased toward "big" languages and needs to learn to respect the nuances of the "small" ones.

In a Nutshell

The paper says: "AI is getting better at translation, but it's still clumsy with regional Chinese languages. We've built a detailed 'error atlas' to show exactly where it trips up, and we've proven that even the smartest AI can't fix these mistakes without human help. Now, we have the map to teach them how to walk properly."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →