← Latest papers
💻 computer science

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

The paper introduces MultiSynt/MT, an open synthetic parallel corpus of 4.8 trillion tokens across 36 European languages generated by translating high-quality English data, which enables multilingual LLMs to achieve superior performance with significantly fewer pre-training tokens compared to native-data baselines while also revealing specific evaluation blind spots in current benchmarks.

Original authors: Maximilian Idahl, Jörg Tiedemann, Sampo Pyysalo, David Salinas, Tomasz Galica, Shenbin Qian, Tudor Nicolae Mateiu, Zihao Li, Anna Lokrantz, Fedor Vitiugin, André F. T. Martins, Jenna Kanerva, Filip Gi
Published 2026-07-02
📖 4 min read☕ Coffee break read

Original authors: Maximilian Idahl, Jörg Tiedemann, Sampo Pyysalo, David Salinas, Tomasz Galica, Shenbin Qian, Tudor Nicolae Mateiu, Zihao Li, Anna Lokrantz, Fedor Vitiugin, André F. T. Martins, Jenna Kanerva, Filip Ginter, Matthias Lindemann, Tim Isbister, Birger Moell, Jonas Lindh, Jan Hajič, Jenia Jitsev, Andrey Kutuzov, Stephan Oepen, Gema Ramírez-Sánchez

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a giant, super-smart robot how to speak and understand 36 different languages. The problem is, the robot's favorite teacher (the internet) mostly speaks English. While there are mountains of English books, articles, and websites, there are only tiny, dusty piles of high-quality text for languages like Maltese, Irish, or Icelandic.

The authors of this paper, MultiSynt/MT, decided to solve this by building a massive "translation factory."

The Big Idea: The Translation Factory

Instead of waiting for people to write 4.8 trillion new words in 36 different languages (which would take forever), they took 100 billion high-quality English documents (the "Nemotron-CC" dataset) and ran them through a super-advanced translation machine.

Think of it like this:

  • The Source: A library of the best English writing available.
  • The Workers: They didn't just use one translator. They used a team of different AI translators (some are like specialized linguists, others are general-purpose smart robots) to translate every single English sentence into 36 target languages.
  • The Result: A new library containing 4.8 trillion tokens (words and symbols) in those 36 languages. For many smaller European languages, this new library is 100 times bigger than any other free library that existed before.

The Experiment: Does "Translated" Work?

The team trained a new AI model using this massive translated library and compared it to a model trained on "native" data (text originally written by humans in those languages).

The Surprise Result:
The model trained on the translated data performed better than the native model, but it got there using 72% less training time and data.

  • Analogy: Imagine two students taking a test. Student A (Native) studied a thick textbook. Student B (Translated) studied a condensed, high-quality summary. Student B not only passed but scored higher, and they did it in less than half the time.

However, the paper warns that this isn't a magic bullet for everything.

The Catch: The "Uncanny Valley" of Language

While the translated data is great for general knowledge and logic, it has a specific weakness: Culture and Idioms.

  • The Metaphor: Imagine a tourist who has studied a phrasebook perfectly. They can order food and ask for directions (general tasks) flawlessly. But if you ask them about a local joke, a specific historical reference, or a slang phrase that only locals use, they might get it wrong because the phrasebook was translated from English, not born from the local culture.
  • The Finding: When the team tested the AI on Norwegian tasks involving local idioms or cultural nuances, the model trained on native human writing still won. The translated model was fluent but lacked the "soul" of the local culture.

The "Blind Spot" in Testing

The paper also discovered a funny flaw in how we usually test AI.

  • The Problem: Most tests are multiple-choice questions (like a standardized exam). These tests are like a "fill-in-the-blank" game. They can't tell the difference between a sentence that sounds perfectly natural and one that sounds slightly "off" or robotic (a phenomenon called "translationese").
  • The Discovery: When the researchers used a different kind of test—asking the AI to write a story and having another AI judge the flow and beauty of the writing—they found that the translated models did have subtle differences. The standard tests were "blind" to these quality gaps, but the "story judge" could see them clearly.

The Conclusion

The paper concludes that translated data is a powerful supplement, not a perfect replacement.

  • For small languages: It's a lifesaver. It provides a massive amount of high-quality data where none existed before.
  • For big languages: It's a great way to fill in the gaps.
  • The Best Strategy: Use the translated data to build a strong foundation, but keep the native human data to teach the AI the cultural nuances, jokes, and local flavor that machines can't invent on their own.

The authors have released this massive "translation library" for free, allowing other researchers to experiment with how to best mix these translated texts with native ones to build the next generation of multilingual AI.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →