← Latest papers
💬 NLP

HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models

The paper introduces HPLT 3.0, an open initiative providing a 30-trillion-token multilingual dataset for nearly 200 languages, accompanied by an open-source processing pipeline, comprehensive evaluation benchmarks for nine European languages, and a suite of pre-trained monolingual and parallel models.

Original authors: Stephan Oepen, Nikolay Arefev, Mikko Aulamo, Marta Bañón, Maja Buljan, Laurie Burchell, Lucas Charpentier, Pinzhen Chen, Mariya Fedorova, Ona de Gibert, Barry Haddow, Jan Hajič, Jindřich Helcl, Andrey
Published 2026-04-21
📖 5 min read🧠 Deep dive

Original authors: Stephan Oepen, Nikolay Arefev, Mikko Aulamo, Marta Bañón, Maja Buljan, Laurie Burchell, Lucas Charpentier, Pinzhen Chen, Mariya Fedorova, Ona de Gibert, Barry Haddow, Jan Hajič, Jindřich Helcl, Andrey Kutuzov, Veronika Laippala, Zihao Li, Risto Luukkonen, Bhavitvya Malik, Vladislav Mikhailov, Amanda Myntti, Dayyán O'Brien, Lucie Poláková, Sampo Pyysalo, Gema Ramírez Sánchez, Janine Siewert, Pavel Stepachev, Jörg Tiedemann, Teemu Vahtola, Dušan Variš, Fedor Vitiugin, Tea Vojtěchová, Jaume Zaragoza

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a giant, super-smart robot how to speak and understand almost every human language on Earth. To do this, the robot needs to "eat" a massive amount of text—books, websites, articles, and conversations. This is called pre-training data.

For a long time, the biggest "buffets" for these robots were mostly filled with English food, with just a few side dishes in other languages. The HPLT 3.0 paper is about a massive new project that has built the world's largest, most diverse, and highest-quality digital library specifically for training these AI robots.

Here is a breakdown of what they did, using some everyday analogies:

1. The Massive Library (The Data)

Think of the internet as a giant, messy attic filled with billions of boxes of old papers, some written in perfect English, others in tiny, rare dialects, and many covered in dust or torn up.

  • The Collection: The HPLT team gathered 30 trillion "tokens" (chunks of words). That's like filling a library with so many books that if you read one every second, it would take you thousands of years to finish.
  • The Diversity: Unlike previous libraries that were 90% English, this one is a true global potluck. About half of the food is in languages other than English, covering nearly 200 different languages and scripts.
  • The Source: They didn't just buy books; they went into the "attic" of the internet (web archives from the Internet Archive and Common Crawl) and pulled out everything they could find.

2. The Cleaning Crew (Data Processing)

You can't just feed a robot raw, dusty attic junk. It would get confused or learn bad habits. The team built a sophisticated assembly line (a pipeline) to clean this data:

  • Text Extraction: They used a smart tool (like a high-tech vacuum) to suck out the actual story from a webpage, ignoring the ads, pop-ups, and "Click Here" buttons.
  • Language ID: They taught a robot to look at a page and say, "Ah, this is Finnish," or "This is actually Spanish, not Portuguese."
  • The "Deduplication" Filter: Imagine finding the same Wikipedia article copied 50 times in your attic. The team removed these copies so the robot doesn't think the same sentence is 50 different facts. They did this globally, ensuring the robot sees unique ideas, not just repeats.
  • Quality Control: They added "stamps" to the documents. Some stamps say "High Quality," others say "Warning: Adult Content" or "Warning: This looks like spam." This lets researchers pick only the best books for the robot to study.

3. The Taste Test (Evaluation)

How do you know the robot is actually learning? You can't just ask it, "Do you feel smart?" You have to give it a test.

  • The Exam: The team created a special exam (HPLT-e) for nine different European languages. It's like a standardized test (think SATs or IELTS) but for AI.
  • The Results: They trained robots on HPLT 3.0 data and compared them to robots trained on other famous datasets. The HPLT 3.0 robots scored higher, proving that cleaner, better-organized data makes smarter robots.
  • The "Quality" Lesson: They also tested if it matters which books you feed the robot. They found that feeding the robot only the "Top 10%" highest-quality documents didn't work as well as feeding it a mix. It's like a student: reading only the best textbooks is good, but reading a variety of good books helps them understand the world better.

4. The Specialized Tools (Models)

The team didn't just give away the library; they also built the robots themselves.

  • Monolingual Models: They built a specific robot for almost 60 different languages. Think of this as hiring a specialist who is a master of only French, or only Swahili, rather than a generalist who knows a little bit of everything. These specialists are incredibly good at understanding the nuances of their specific language.
  • Translation: They also built a massive collection of sentences where one side is in Language A and the other is in Language B. This is like a giant bilingual dictionary that helps robots learn to translate instantly.

5. The "Synthetic" Kitchen (Machine Translation)

For some very rare languages, there just aren't enough books in the attic. So, the team used a clever trick:

  • They took high-quality English texts (like educational articles), translated them into rare languages using AI, and added them to the library.
  • The Analogy: It's like taking a perfect recipe written in English, translating it into a rare dialect, and giving it to a chef who has never seen that dish before. It helps the chef learn the basics of the cuisine, even if they haven't eaten the real thing yet.

6. The Ethical Safety Net

The paper is very careful about ethics.

  • Privacy: They scrubbed the data to remove personal info like phone numbers and email addresses.
  • Copyright: They respected "Do Not Enter" signs (robots.txt) on websites and removed content that was clearly copyrighted or prohibited.
  • Transparency: Unlike some big tech companies that keep their secret recipes, HPLT is an open-source project. They are sharing the library, the cleaning tools, and the robots for free, so anyone can use them.

The Big Picture

The HPLT 3.0 project is a massive effort to democratize AI. Instead of only big corporations having the fuel to build super-smart AI, this project is handing out high-quality fuel to researchers, universities, and developers everywhere.

Their goal is to ensure that the future of AI isn't just an "English-speaking" robot, but a truly multilingual one that respects and understands the diversity of human culture. They are essentially building the "public library" for the next generation of artificial intelligence.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →