← Latest papers
💬 NLP

EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training

This paper demonstrates that combining continued pretraining with Estonian-enriched multilingual replay and lightweight post-training alignment substantially enhances Estonian language capabilities in multilingual LLMs while effectively preserving their English performance and general reasoning skills.

Original authors: Aleksei Dorkin, Taido Purason, Emil Kalbaliyev, Hele-Andra Kuulmets, Marii Ojastu, Mark Fišel, Tanel Alumäe, Eleri Aedmaa, Krister Kruusmaa, Kairit Sirts

Published 2026-08-26
📖 5 min read🧠 Deep dive

Original authors: Aleksei Dorkin, Taido Purason, Emil Kalbaliyev, Hele-Andra Kuulmets, Marii Ojastu, Mark Fišel, Tanel Alumäe, Eleri Aedmaa, Krister Kruusmaa, Kairit Sirts

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Large language models are the engines behind many of today's most advanced artificial intelligence tools, capable of writing stories, solving problems, and answering questions. These systems are trained on vast amounts of text from the internet, learning patterns of human language by reading billions of examples. However, the data used to teach them is heavily skewed toward English. Because English dominates the digital world, these models become experts in English while often struggling with smaller languages like Estonian, which have far less digital representation. This imbalance means that while an AI might write a perfect essay in English, it could produce confusing or incorrect text when asked to do the same in Estonian. Researchers are now working to fix this gap, asking whether it is possible to teach these powerful, English-heavy models a new language without breaking the skills they already possess.

A team of researchers from Estonia set out to answer this question by creating a specialized version of a large language model tailored for the Estonian language. They started with two existing models: one that was heavily trained on English data and another that was already somewhat familiar with many languages. The researchers did not build a new model from scratch. Instead, they took these existing systems and gave them a focused "re-education" using a specific mix of data. First, they exposed the models to a large volume of Estonian text, including literature, academic writing, and web content, to build a foundation of language knowledge. To ensure the models did not forget how to reason or understand other topics, they mixed in English text, code, and mathematical problems during this training phase. This process, known as continued pretraining, allowed the models to learn Estonian while keeping their general intelligence intact.

After this initial training, the researchers refined the models to make them better at following instructions, a crucial skill for chatbots and assistants. They taught the models how to respond to user commands in both Estonian and English, using a mix of real human conversations and synthetic examples. A key step in their process involved a technique called chat vector merging. Imagine a model as a student who has learned a new language but has forgotten how to speak politely or follow complex rules. The researchers took the "knowledge" of how to follow instructions from the original English version of the model and blended it into their new Estonian-trained version. This allowed the final product to speak fluent Estonian while retaining the ability to follow instructions clearly in both languages.

The results of this experiment were surprising and significant. Before the training began, the model that was already multilingual performed better on Estonian tasks than the one that was primarily English-focused. However, after the adaptation process, the situation flipped. The model that started with a strong English foundation showed much larger improvements in Estonian, eventually outperforming the multilingual model. This suggests that a model's starting point in a specific language does not necessarily predict how well it will learn that language later. The researchers found that the adapted models could understand grammar, translate text, and solve logic puzzles in Estonian far better than before. They also tested these models in a public arena where human users could chat with them and vote for the best response. The new Estonian models ranked highly, competing successfully against much larger and more complex systems.

The study also revealed important limits to how these models learn. The researchers tested whether repeating the Estonian training data multiple times would help the models learn better, a method some other studies have used. They found that for this size of model, going through the data a second time offered no real benefit and was simply a waste of computing power. One training cycle was sufficient. Furthermore, the researchers discovered that while the models became excellent at Estonian, there was a slight trade-off: their ability to follow instructions in English dipped slightly. However, the chat vector merging technique successfully restored most of this lost English ability, proving that it is possible to specialize a model for a smaller language without sacrificing its general capabilities.

This work provides a clear path forward for improving artificial intelligence in languages that are often overlooked. By carefully mixing new language data with existing knowledge and using smart techniques to blend different model behaviors, researchers can create powerful tools for specific communities without needing to start from zero. The findings suggest that the most effective way to adapt a large language model might not be to start with a model that already knows the language, but to take a model that is strong in general reasoning and teach it the new language with a balanced diet of data. The resulting models are now available for others to use, offering a high-quality, native-level experience for Estonian speakers while maintaining the robustness of a global intelligence.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →