← Latest papers
💬 NLP

GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training

The paper introduces the GPT-NL Public Corpus, a permissively licensed dataset featuring 36 billion unique Dutch tokens alongside curated English, code, and German/Danish data, designed to facilitate the development of lawful and useful commercial language models.

Original authors: Jesse van Oort, Frank Brinkkemper, Erik de Graaf, Bram Vanroy, Saskia Lensink

Published 2026-04-03
📖 5 min read🧠 Deep dive

Original authors: Jesse van Oort, Frank Brinkkemper, Erik de Graaf, Bram Vanroy, Saskia Lensink

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to teach a brilliant, but very young, child (an Artificial Intelligence) how to speak, think, and write in Dutch. To do this, you need to give them a massive library of books, newspapers, and conversations to read.

However, there's a catch: You can't just grab any book off the shelf. You have to make sure you have the legal right to photocopy it, and you don't want to give the child books filled with hate speech, lies, or confusing nonsense.

This paper is the story of how a team in the Netherlands built a special, super-legal, and super-clean library called the GPT-NL Public Corpus specifically for teaching AI to speak Dutch.

Here is the breakdown of their project using simple analogies:

1. The Problem: The "Empty Bookshelf"

For a long time, if you wanted to train an AI to speak English, you could dump the entire internet into its brain. But for Dutch? The "bookshelf" was mostly empty, or the books were locked behind paywalls, or they were written in a way that made them illegal to use for commercial AI.

Because of this, Dutch speakers had to rely on "multilingual" AIs that were mostly trained on English. It's like trying to learn to play the piano by only listening to a violinist; you get the music, but the technique is all wrong.

2. The Solution: The "GPT-NL Public Corpus"

The team built a new library. It's not just a pile of random web pages; it's a curated collection of 21 specific Dutch-language collections.

  • The Size: It contains about 36 billion words (tokens) of pure Dutch that no other AI has seen before.
  • The Bonus: They also added some English, German, Danish, and Code (programming language) to help the AI learn better, just like a child learns better if they hear a few other languages nearby.

3. The Three Golden Rules (The "Safety Filters")

Before any book could enter this library, it had to pass three strict tests:

  • Rule #1: Is it Useful?
    • Analogy: We don't want the AI to memorize random spam emails or nonsense. We want it to learn facts, logic, and clear communication. They focused on data that helps the AI be a helpful assistant, not a hallucinating dreamer.
  • Rule #2: Is it Legal?
    • Analogy: Imagine a "No Trespassing" sign on a house. The team only picked books where the owner said, "Yes, you can copy this!" (Creative Commons licenses). They avoided books with "No Commercial Use" signs or "No Derivatives" signs. They wanted to build a library that anyone, even a business, could legally use.
  • Rule #3: Is it Safe?
    • Analogy: If you give a child a book full of bullying or hate speech, they might start acting that way. The team scrubbed out toxic content, bias, and offensive language. They wanted the AI to be polite and helpful.

4. How They Filled the Shelves (The Sources)

Since they couldn't just scrape the whole internet, they had to get creative:

  • The "Time Travelers" (Archives): They partnered with Dutch archives to digitize old books and newspapers (over 100 years old). Since these are so old, they are in the "Public Domain" (free for everyone to use).
  • The "Government Friends": They worked with the Dutch government to get official documents, court rulings, and parliamentary records. These are usually public, but hard to find in a clean format.
  • The "Synthetic Builders" (Type 1 Only): They didn't just ask an AI to "make up" new stories (which can be fake). Instead, they took fact databases (like Wikidata) and used a smart script to turn dry facts into readable sentences. Think of it like turning a spreadsheet of "Name: Willem-Alexander, Title: King" into a sentence: "Willem-Alexander is the King."
  • The "Translator" (YouTube): They took English and Spanish YouTube transcripts, cleaned them up, and used a translation tool to turn them into Dutch. This gave the AI a chance to hear how real people speak.
  • The "Web Sweeper" (C5): They used a special tool to scan the web, but only for pages that explicitly said, "This content is free to use." They filtered out everything else to be 100% sure about the legal rights.

5. The Quality Control (The "Librarian's Inspection")

Before the books went on the shelf, a team of human librarians (data curators) did a final check.

  • They looked at samples to make sure the Dutch sounded natural.
  • They checked for personal secrets (like phone numbers or addresses) and scrubbed them out.
  • They rated the "risk" of each book. If a book had a little bit of bias, they might use less of it when training the AI, just to be safe.

6. Why This Matters

This paper isn't just about a dataset; it's about fairness and safety.

  • For Researchers: It gives them a clean, legal starting point to build better Dutch AIs without worrying about getting sued.
  • For the Public: It ensures that the AI tools we use in the future will understand Dutch culture, laws, and language better, without being poisoned by bad data or illegal content.

In a nutshell: The GPT-NL team didn't just dump a bucket of water (data) into a pool. They built a filtered, legal, and high-quality swimming pool specifically for Dutch-speaking AI, ensuring it's safe for everyone to dive in.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →