← Latest papers
💬 NLP

LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering

This paper demonstrates that parameter-efficient Spoken Question Answering for the low-resource language Luxembourgish can be effectively achieved by training on synthetic speech generated from translated text QA pairs using multiple TTS systems, revealing that task-specific performance depends more on diverse voice configurations than on standard TTS quality metrics.

Original authors: Nina Hosseini-Kivanani, Marco Matassoni, Alessio Brutti

Published 2026-07-07
📖 4 min read☕ Coffee break read

Original authors: Nina Hosseini-Kivanani, Marco Matassoni, Alessio Brutti

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to build a smart assistant that can answer questions spoken in Luxembourgish, a language spoken by only a few hundred thousand people. The problem? To teach a computer to understand spoken language, you usually need thousands of hours of real people recording questions and answers. But for Luxembourgish, that data simply doesn't exist yet.

This paper, LuxSQA, asks a clever question: Can we cheat by using "fake" voices (Text-to-Speech or TTS) to teach the computer instead?

Here is the story of how they did it, explained simply:

1. The Problem: The "Empty Library"

Think of training a smart AI like teaching a student. If you want to teach them Luxembourgish, you need a library of books (text) and recordings (audio).

  • The Text Library: Full of questions and answers.
  • The Audio Library: Empty. There are no recordings of people asking these questions.

Usually, you'd have to hire actors to record thousands of hours of audio. That's expensive and slow.

2. The Solution: The "Robot Voice Factory"

Instead of hiring actors, the researchers built a Robot Voice Factory.

  • Step 1: They took existing text questions (from English) and translated them into Luxembourgish.
  • Step 2: They used different AI "Voice Actors" (TTS systems) to read these questions out loud.
  • Step 3: They paired these robot voices with the correct text answers.

Now, they had a massive "fake" library of spoken questions to train their AI.

3. The Experiment: Testing Different "Voice Actors"

The researchers didn't just use one robot voice. They tried five different types of TTS engines (like MMS-TTS, Qwen, and OmniVoice). Some were designed to clone a specific person's voice, while others were designed to create a brand-new, unique voice from scratch.

They also tried two main strategies:

  • The Single Voice: Training the AI on questions spoken by just one type of robot voice.
  • The Mixed Choir: Training the AI on a huge mix of questions spoken by many different robot voices (about 230,000 questions in total).

4. The Surprising Discovery: "Good Sounding" isn't "Good Learning"

Here is the most interesting part. The researchers had a "Sound Quality Meter" (a tool that scores how natural a robot voice sounds to a human ear).

  • Expectation: They thought the robot voices that sounded the most natural and human-like would make the best teachers.
  • Reality: The voices that sounded the best to human ears actually made the AI perform worse at answering questions.
  • The Winner: The AI learned best when it was trained on a diverse mix of voices, even if some of those voices sounded a bit more robotic or artificial.

The Analogy: Imagine learning to recognize a friend's face. If you only look at one perfect, high-definition photo, you might get confused when they wear a hat or stand in the dark. But if you look at a messy pile of photos—some blurry, some in black and white, some with different lighting—you learn to recognize the person much better. The AI needed the "messy pile" of diverse voices to learn the language, not just the "perfect" voice.

5. The Result: A Working System Without Real Humans

By using this "Mixed Choir" of robot voices, the researchers built a system that could listen to a spoken Luxembourgish question and give a text answer.

  • When they tested it on real human speakers (two different people), the system performed surprisingly well.
  • It proved that you don't need a massive library of real human recordings to build a spoken question-answering system for rare languages. You can build a very capable one using synthetic data, as long as you use the right kind of synthetic data.

The Bottom Line

This paper shows that for languages with few speakers, we don't need to wait for thousands of people to record themselves. We can use AI to generate the training data, but we must be careful: Quantity and variety matter more than perfection. A diverse mix of "imperfect" robot voices teaches the AI better than a single "perfect" robot voice.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →