← Latest papers
💻 computer science

Luth: Efficient French Specialization for Small Language Models and Cross-Lingual Transfer

This paper introduces Luth, a family of French-specialized Small Language Models that achieve state-of-the-art performance on French benchmarks while retaining English capabilities through targeted post-training and strategic model merging, effectively addressing the performance gap in non-English SLMs.

Original authors: Maxence Lasbordes, Sinoué Gad

Published 2026-02-10
📖 4 min read☕ Coffee break read

Original authors: Maxence Lasbordes, Sinoué Gad

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of Artificial Intelligence (AI) as a massive library. For a long time, almost all the books in this library have been written in English. If you ask the librarian (the AI) a question in French, they might stumble, give a vague answer, or sound like they are reading a dictionary rather than having a conversation. This is especially true for the "smaller" librarians (Small Language Models or SLMs), who are designed to be fast and efficient but often lack deep knowledge of languages other than English.

The paper you shared introduces a new team of librarians called Luth. Their goal was simple: take these efficient, small AI models and teach them to speak French fluently without making them forget how to speak English.

Here is how they did it, using some everyday analogies:

1. The Problem: The "English-Only" Bias

Most AI models are trained on a diet of mostly English data. It's like teaching a student only in English; they might be great at English literature but struggle with French history or math problems written in French. The authors noticed that even the best "multilingual" models were still significantly weaker in French compared to English.

2. The Solution: A Specialized "French Boot Camp" (Luth-SFT)

To fix this, the authors didn't just throw more random French text at the AI. Instead, they built a specialized training camp called Luth-SFT.

  • The Curriculum: They gathered 570,000 high-quality "instruction and answer" pairs. Think of this as a massive workbook of French questions and perfect answers.
  • The Translation Trick: They took excellent English textbooks (datasets) and didn't just translate the answers word-for-word. Instead, they translated the questions into French and then asked the AI to write new answers from scratch. This ensured the French sounded natural and human, not robotic.
  • The "Scholar" Section: A special part of this workbook focused on hard academic subjects (like Math, Physics, and Engineering) using real exam papers from French high schools and universities. This gave the AI a serious "brain boost" in reasoning and logic.

3. The Training: Full Fine-Tuning

They took existing small AI models (the "base" models) and ran them through this French boot camp. This is like taking a generalist athlete and putting them through a rigorous, specialized training regimen to become a French-speaking expert.

The Catch: When you train a model intensely on one language, it sometimes starts to forget its other skills. In this case, the models got great at French but started to get slightly worse at English.

4. The Magic Trick: "Model Merging" (The Smoothie Effect)

This is where the paper gets clever. To fix the "forgetting English" problem without losing the new French skills, they used a technique called Model Merging.

Imagine you have two smoothies:

  1. Smoothie A: The original model (Great at English, okay at French).
  2. Smoothie B: The newly trained model (Amazing at French, slightly weaker at English).

Instead of choosing one or the other, they "blended" them together. They mixed the "brain" of the original model with the "brain" of the trained model.

  • The Result: The new blended model (Luth) kept the French superpowers and regained (or even improved) its English skills. It was like getting the best of both worlds in a single cup.

5. The Results: The New Champions

The authors tested these new Luth models against other small AI models on six different challenges (like math problems, following complex instructions, and general knowledge).

  • In French: The Luth models crushed the competition. They outperformed every other open-source model of the same size, improving scores by up to 11% on average.
  • In English: Surprisingly, they didn't just stay the same; they actually got better at English too. The authors call this "cross-lingual transfer," suggesting that learning French deeply actually helped the model understand English concepts better as well.

Summary

The paper claims that by creating a high-quality French dataset and using a "blending" technique to merge models, they created a family of small, efficient AI models that are now the best at French available, without sacrificing their English abilities. They proved you don't need a giant, expensive supercomputer to make a great French AI; you just need the right data and the right mixing method.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →