← Latest papers
💬 NLP

Training Models on Dialects of Translationese Shows How Lexical Diversity and Source-Target Syntactic Similarity Shape Learning

This paper demonstrates that training small English language models on machine-translated data reveals that while general perplexity is primarily driven by the lexical diversity of the translated corpus, grammatical performance is strongly shaped by the typological similarity between the source language and English.

Original authors: Jenny Kunz

Published 2026-02-19
📖 5 min read🧠 Deep dive

Original authors: Jenny Kunz

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to teach a robot how to speak perfect English. Ideally, you'd feed it millions of books, news articles, and conversations written by native English speakers. But what if you don't have enough native text? What if you have to use books that were originally written in French, Japanese, or Swahili and then translated into English?

This is the problem Jenny Kunz tackled in her paper. She asked: Does it matter which language we translate from? And does the "flavor" of that translation change how the robot learns?

Here is the story of her findings, explained with some everyday analogies.

The "Accent" of Translation

First, the paper introduces a concept called Translationese. Think of native English text as a clear, crisp recording of a singer. Translationese is like a cover song. It's the same song, but the singer has a slight accent, maybe they sing the notes a little differently, or they skip a few complex harmonies.

Even professional human translators (and definitely AI translators) leave traces of their original language behind. A sentence translated from German might sound a bit stiff or structured differently than one written by a native speaker.

The Experiment: 24 Different Accents

Kunz took 24 different languages (ranging from Germanic cousins like Swedish to distant relatives like Swahili and Tamil) and translated huge amounts of text into English. She then trained small AI models on these "translated English" datasets.

She wanted to see two things:

  1. How well does the robot predict the next word? (Like a autocomplete feature).
  2. Does the robot actually understand grammar? (Like knowing the difference between "The cat is sleeping" and "The cat are sleeping").

The Big Discoveries

1. The "Vocabulary Buffet" vs. The "Grammar Gym"

The paper found that the type of language you translate from affects the robot differently depending on what you are testing.

  • For General Fluency (The Vocabulary Buffet):
    When the robot is just trying to guess the next word in a sentence, the variety of words in the translated text matters most.

    • Analogy: Imagine you are learning to cook. If you are given a cookbook translated from a language with a huge, diverse vocabulary (like a massive buffet), your robot learns to predict ingredients well, even if the grammar is a bit weird. It doesn't matter if the book came from a language similar to English; what matters is that the book had lots of different words to learn from.
    • Result: In small data sets, the "richness" of the translated text was the biggest predictor of success.
  • For Grammar (The Grammar Gym):
    When the robot needs to learn complex rules (like subject-verb agreement or long-distance connections in a sentence), how similar the source language is to English becomes the hero.

    • Analogy: Imagine teaching someone to play tennis. If you teach them using a racket that looks and feels exactly like a standard tennis racket (a language structurally similar to English, like Swedish), they learn the proper swing faster. If you teach them with a racket that feels totally different (a distant language like Finnish), they struggle to learn the specific mechanics, even if you give them more practice time.
    • Result: Once the robot had enough data, models trained on languages similar to English (like Swedish or Dutch) became much better at grammar than those trained on distant languages.

2. The "Flattening" Effect

No matter which language was used, the translated text always made the robot slightly "worse" than if it had learned from native English.

  • Analogy: It's like listening to a song played on a cheap speaker versus a high-end stereo. The cheap speaker (translated text) flattens the sound. It loses some of the nuance, the "sparkle," and the natural flow of the original. The robot trained on translations is always a bit less "native" than one trained on real English.

3. The "Family Reunion" Effect

The most fascinating finding was about how these robots interact with each other.

  • Analogy: Imagine two people who learned English by translating from two different languages. If those two source languages are "cousins" (structurally similar), the two robots understand each other's "accents" perfectly. If the source languages are strangers, the robots get confused.
  • Result: A robot trained on Swedish-translated text could easily understand text translated from Norwegian. But it struggled with text translated from a distant language like Arabic. This proves that translation creates specific "dialects" of English, and those dialects are shaped by the original language.

The Bottom Line

If you are building an AI and you have to use translated data:

  1. If you just want the AI to sound okay generally: Pick a source language that has a rich, diverse vocabulary.
  2. If you need the AI to be grammatically perfect: Pick a source language that is structurally similar to English (like other European languages).
  3. The Reality Check: Translated data is a great backup plan when native text is scarce, but it will always leave the AI with a slight "accent" and a slightly weaker grasp of complex grammar compared to a native speaker.

In short: Translation is a bridge, but it's not the same as walking on the native soil. The closer the bridge is to the native soil (typological similarity), the better the robot learns the rules of the road.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →