LuxMT Technical Report
This paper introduces LuxMT, a machine translation system fine-tuned on Gemma 3 27B for Luxembourgish-to-French and English translation using data from LuxAlign and Luci, which demonstrates significant performance improvements over the baseline and explores the potential of LuxEmbedder as a quality estimation metric.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, multilingual robot named Gemma. This robot has read almost everything on the internet and speaks many languages fluently. However, there's a problem: while Gemma knows a lot about French, English, and German, it's a bit clumsy when it comes to Luxembourgish (a small, local language spoken in Luxembourg). It often stumbles, misses nuances, or sounds like a textbook rather than a local.
The author of this paper, Nils, decided to give Gemma a specific "boot camp" to fix this. The result is LUXMT, a super-charged version of Gemma designed specifically to translate Luxembourgish into French and English.
Here is the story of how they built it, broken down into simple steps:
1. The Problem: "Don't Cheat on the Test!"
Before training the robot, Nils needed to know if Gemma was actually good at Luxembourgish or just guessing.
- The Trap: Usually, scientists use standard test questions (like FLORES-200) to grade robots. But Nils worried that the robot had already seen these questions while it was learning on the internet. That would be like a student memorizing the answer key before taking the exam.
- The Solution: Nils created a brand new test using articles from Luci, a tourist magazine about Luxembourg. Since this magazine is written specifically for visitors and translated by humans, it's a fresh, fair test that the robot hasn't seen before.
2. The Training Data: Cleaning the "Library"
To teach the robot, Nils needed a library of books where the Luxembourgish text sat right next to its French or English translation.
- The Source: He used news articles from RTL (a local radio/TV station) and transcripts from the Luxembourgish parliament.
- The Filter (LUXEMBEDDER): Here's the tricky part. Not every news article is a perfect 1-to-1 translation. Sometimes a sentence in Luxembourgish doesn't match the French sentence next to it perfectly.
- Nils used a special tool called LUXEMBEDDER to act like a strict librarian. This tool reads both sentences and calculates how similar they are. If the librarian thinks, "These two sentences don't really mean the same thing," it throws the pair in the trash.
- The Result: They threw away a lot of "noisy" data, keeping only the cleanest, most accurate translations to teach the robot.
3. The Training: "One Day is Enough"
Nils tried teaching the robot for different amounts of time (epochs).
- He tried 1 day, 2 days, and 3 days of training.
- The Surprise: The robot learned the most after just one day of intense training. Training it longer actually made it slightly worse, like over-practicing a song until you forget the lyrics. So, they stopped after one round.
4. The Results: A Happy Accident
When they put the trained robot (LUXMT) to the test:
- French & English: It got much better at translating Luxembourgish into these languages, beating the original "untuned" Gemma by a huge margin.
- The German Surprise: Even though they never showed the robot any German data during training, it got better at translating Luxembourgish to German too!
- The Analogy: Imagine teaching someone to speak Italian and Spanish. Suddenly, they start speaking Portuguese better, even though you never taught them Portuguese. This happens because the languages are cousins; learning one helps you understand the others.
5. The New Tool: The "Quality Checker"
Nils also tested his "Librarian" tool (LUXEMBEDDER) to see if it could act as a judge.
- Usually, to grade a translation, you need a human to compare the robot's work against a perfect human translation.
- Nils wondered: Can the Librarian tool just look at the robot's output and say, "This looks good" without needing a human reference?
- The Verdict: It seems to work pretty well and agrees with other computer judges. However, Nils warns us to be careful. It's a promising new tool, but like a new driver, it needs more practice before we trust it with our lives.
Summary
Nils took a smart but slightly clumsy robot, gave it a clean, high-quality diet of Luxembourgish news and parliament speeches, and taught it for just one day. The result is LUXMT, a translator that now speaks Luxembourgish with much more confidence. Along the way, they built a new, fair test to grade future robots and discovered a new way to check translation quality without needing a human to do the grading every time.
The Bottom Line: By being careful with the data and using smart filters, they turned a general-purpose AI into a specialist for a small, unique language.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.