← Latest papers
💬 NLP

Open Machine Translation for Esperanto

This paper presents the first comprehensive evaluation of open-source machine translation systems for Esperanto, demonstrating that the NLLB family outperforms other models across multiple language pairs while releasing the code and best-performing models to the public.

Original authors: Ona de Gibert, Lluís de Gibert

Published 2026-04-01
📖 4 min read☕ Coffee break read

Original authors: Ona de Gibert, Lluís de Gibert

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of languages as a giant, bustling library. Most languages, like English or Spanish, are like massive, well-stocked sections with millions of books, expert librarians, and high-tech scanners. Then, there's Esperanto.

Esperanto is a "constructed language"—it wasn't born from centuries of natural evolution but was built by a doctor named Zamenhof in 1887 like a custom-built tool. Its goal? To be a fair, easy-to-learn bridge between all people. It has a very regular grammar (no tricky exceptions!) and a logical way of building words, kind of like Lego bricks snapping together perfectly.

Despite having a passionate global community and a huge online presence, Esperanto has been largely ignored by the high-tech world of Machine Translation (MT). It's like having a beautiful, functional car in a garage, but the mechanics have never built a specialized engine for it, so it's been running on a generic, clunky engine that doesn't quite fit.

This paper is the story of a team of mechanics (the researchers) who decided to fix that. They asked: "What happens if we try to drive Esperanto using the best engines available today?"

The Race: Who Can Translate Esperanto Best?

The researchers set up a race with three different types of "engines" (translation models) to see which one could translate Esperanto to and from English, Spanish, and Catalan most accurately.

  1. The Rule-Based Engine (Apertium): This is like an old-school dictionary and a strict set of grammar rules. It's fast and lightweight, but it's rigid. If the sentence structure is slightly unusual, it gets confused.
  2. The "Big Brain" Generalists (LLMs like Llama): These are the massive, all-knowing AI models trained on the entire internet. They are like brilliant students who have read every book in the library. They know Esperanto exists, but they haven't practiced translating it specifically.
  3. The Specialized Translators (NLLB and Custom Models): These are models specifically trained to translate many languages. Think of them as professional translators who have spent their whole lives studying the nuances of 200+ languages.

The Results:

  • The Winner: The NLLB family (the specialized translators) won the race by a clear margin. They were the most accurate and fluent.
  • The Surprise: The massive "Big Brain" models (LLMs) did okay, but they weren't as good as the specialized translators. In fact, some of the "translation-specialized" AI models actually performed worse than the general ones, likely because they were over-trained on other languages and forgot how to handle the unique logic of Esperanto.
  • The Underdog: The researchers also built their own tiny, custom engines (small models). Even though they were tiny (like a scooter compared to a truck), they performed surprisingly well, almost as good as the massive ones. This is great news because tiny models are cheap, fast, and can run on a regular laptop.

The Human Test: Does it Sound Real?

Computers use math to grade translations, but humans use their ears and hearts. The researchers asked native speakers to pick the best translation from the top three models.

  • The Verdict: The specialized NLLB model was the crowd favorite, winning about half the time.
  • The Catch: Even the winner wasn't perfect. Sometimes it translated words too literally (like translating "snowboard" as "snow table" because the word is built from "snow" + "board"). Sometimes the AI made up facts or missed details. It's like a very good student who knows the vocabulary but occasionally misunderstands the joke.

Why This Matters: The "Open Source" Spirit

Esperanto isn't just a language; it's a movement built on openness, equality, and community. The researchers felt that relying on secret, corporate AI tools went against the spirit of Esperanto.

So, they did something rare in the tech world: They gave everything away.

  • They released their code.
  • They released their best models for free.
  • They showed that you don't need a billion-dollar supercomputer to translate a language; you just need the right, efficient tools.

The Big Takeaway

This paper proves that even for a "low-resource" language like Esperanto, we can build excellent, open, and efficient translation tools.

  • Don't just use the biggest AI: Sometimes, a specialized, smaller tool works better than a giant, general-purpose one.
  • Community matters: By sharing their work, the researchers are helping the Esperanto community keep their digital infrastructure independent and accessible.
  • The future is bright: With these new tools, Esperanto speakers can communicate with the world more easily, keeping the dream of Zamenhof alive in the digital age.

In short: The researchers took a neglected language, tested the best tools available, found the winners, built their own efficient tools, and handed the keys to the community so everyone can drive forward.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →