Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus
This paper presents a unified pipeline using three state-of-the-art TTS architectures to synthesize high-quality bilingual Quechua and Spanish speech for the Peruvian Constitution, leveraging cross-lingual transfer to overcome data scarcity in the low-resource Quechua language while providing open-source resources for inclusive legal speech technologies.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very important rulebook for a country—the Constitution. In Peru, this rulebook is written in Spanish, which most people understand. But there's a huge group of people who speak Quechua, an ancient and beautiful indigenous language. For them, reading the law is like trying to read a map written in a language you don't know; the information is there, but it's locked behind a wall.
This paper is about building a bridge to let that information flow freely. The authors built a special "voice machine" (Text-to-Speech) that can read the Peruvian Constitution out loud in both Spanish and Quechua, making the law accessible to everyone.
Here is how they did it, explained with some everyday analogies:
1. The Problem: The "Empty Library"
Think of the Spanish language as a massive, fully stocked library with millions of books, audiobooks, and recordings. Now, imagine Quechua is like a tiny, cozy shed with only a few books and very few recordings.
To teach a computer to speak Quechua well, you usually need a huge library of recordings (thousands of hours). But for Quechua, that library is almost empty. If you try to teach a computer with just a few hours of audio, it sounds like a robot with a broken voice box.
2. The Solution: The "Bilingual Tutor"
The authors had a clever idea: Why not teach the computer using the big library (Spanish) to help it learn the small language (Quechua)?
They treated Spanish as a "bilingual tutor." Since Spanish and Quechua share some sounds and rhythms (like how a guitar and a violin share musical notes), the computer could learn the musicality of speech from the abundant Spanish data and apply those skills to the scarce Quechua data.
They used three different types of "voice engines" (AI models) to test this:
- XTTS v2: A powerful, heavy-duty engine.
- F5-TTS: A modern engine that uses "flow" to smooth out the voice.
- DiFlow-TTS: A compact, efficient engine.
3. The Training: Cleaning the Data
Before teaching the computer, they had to clean up the "textbooks."
- The Filter: They had hours of Quechua recordings, but many were just 1-second clips of silence or noise. Imagine trying to learn to sing by listening to 1-second snippets of a song; it doesn't work. They filtered out the short, bad clips and kept only the clear, long sentences.
- The Translator: Quechua is a "sticky" language (agglutinative), meaning words can get very long and complex. They used a digital tool to break these words down into their standard parts, ensuring the computer didn't get confused by spelling variations.
4. The Results: The "Small Engine" Wins
Usually, in the world of AI, bigger is better. You'd think the biggest, most expensive computer model would win. But here, the results were surprising:
- The Heavyweight (XTTS): It was big and powerful, but it sounded a bit stiff.
- The Middleweight (F5-TTS): It sounded very consistent, like a reliable radio host.
- The Lightweight (DiFlow-TTS): This was the surprise winner! Even though it was the smallest model (the "compact engine"), it produced the most natural-sounding Quechua.
The Analogy: Think of it like cooking. You can have a giant, industrial kitchen (the big model), but if you don't have the right ingredients (data), the food tastes bland. The small kitchen (DiFlow) used the "secret sauce" of cross-language learning perfectly, making a delicious meal with very few ingredients.
5. Why This Matters
This isn't just about making a cool voice app. It's about justice and inclusion.
- Accessibility: Now, a Quechua speaker can listen to the Constitution on the radio or a phone, understanding their rights in their own language.
- Reusability: The authors didn't just build the voice; they released the "recipe" (the code and data). Other researchers can use this to build better tools for Quechua, like voice assistants or translation apps.
The Bottom Line
The authors proved that you don't need a massive amount of data to teach a computer a rare language. If you are smart about how you connect it to a related, well-resourced language (like Spanish), you can build a high-quality voice that respects and elevates indigenous cultures.
They effectively gave a voice to the Constitution, ensuring that the law speaks to everyone, not just those who speak the dominant language.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.