← Latest papers
💬 NLP

Building Community-Centred NLP Resources for Puno Quechua

This paper introduces the first dedicated automatic speech recognition resources for Puno Quechua, comprising a 66-hour participatory speech corpus, a systematic benchmark of state-of-the-art models, and the open release of all datasets and fine-tuned models to support community-centered language preservation.

Original authors: Elwin Huaman, Adrian Gamarra Lafuente, Johanna Cordova, Anna Korhonen

Published 2026-05-28
📖 5 min read🧠 Deep dive

Original authors: Elwin Huaman, Adrian Gamarra Lafuente, Johanna Cordova, Anna Korhonen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine language as a vast, ancient library. For centuries, the books in the "Puno Quechua" section of this library have been left in the dark, with no digital catalog or search engine to help people find them. While the world rushes to build AI tools that speak English, Spanish, or Mandarin, the 465,000 people who speak Puno Quechua in Peru have been locked out of the digital conversation because they can't type in their own language.

This paper is like a team of librarians and engineers building the first-ever digital key to unlock that section. They didn't just build a key; they built a whole new system designed specifically for the people who live there.

Here is the story of what they did, broken down into simple parts:

1. The Problem: A Language Without a Voice in the Digital World

Think of AI tools like ChatGPT as a very strict librarian who only speaks a few major languages. If you try to talk to them in Puno Quechua, they just stare blankly. Why? Because the "training data" (the books the librarian read) for Puno Quechua was missing. Previous attempts to teach computers Quechua were like trying to teach a dog to speak by mixing up instructions for a cat, a hamster, and a dog. They treated all Quechua dialects as one big, identical group, ignoring the unique sounds and rules of Puno Quechua.

2. The Solution: Building a Library with the Community

Instead of just grabbing data from the internet, the researchers went into the Puno region and built a community garden. They used a "participatory design" approach, which is like inviting the neighbors to help plant the seeds rather than just handing them a pre-made garden.

  • The Harvest: They gathered 66 hours of voice recordings. Imagine 66 hours of people reading stories, chatting about their day, and talking about farming, health, and technology.
  • The Quality Control: Out of this, 36 hours were carefully checked and transcribed by hand by native speakers (the "gold standard"). The rest was a mix of spontaneous speech and "silver" data (automatically transcribed but less perfect, like a rough draft).
  • The Result: This is now the largest collection of Puno Quechua speech data ever created for a single variety of the language.

3. The Experiment: Teaching the AI to Listen

Once they had the "books" (the data), they had to teach the AI how to read them. They tried three different "students" (AI models):

  • Whisper: A generalist student who knows many languages.
  • wav2vec2: A student who learned mostly from English books.
  • XLS-R: A super-student who has read books in 128 different languages.

They tested these students in two ways:

  1. Reading Aloud: Testing them on clear, scripted sentences.
  2. Eavesdropping: Testing them on spontaneous, real-life conversations (which are messy and fast).

They also tried a special trick called Continued Pre-Training (CPT). Imagine taking the "super-student" (XLS-R) and giving them a summer camp where they just listened to Puno Quechua radio and conversations without needing to read the words first. This helped their ears get used to the unique sounds of the language (like the sharp "ejective" sounds) before they started the hard work of learning to read.

4. The Results: What Worked Best?

The team found some surprising things, like a coach discovering which training drills actually help the team win:

  • The "Silver" Secret: For messy, real-life conversations, the "rough draft" data (silver transcriptions) was the secret weapon. Even though it wasn't perfect, adding it to the training data reduced errors by 77%. It's like giving a student a messy notebook full of ideas; it helps them understand the flow of conversation better than just having perfect, clean notes.
  • The Summer Camp Effect: The "Continued Pre-Training" (CPT) helped the AI understand the unique sounds of Puno Quechua much better, especially for clear, scripted speech. It lowered errors by 42% compared to skipping this step.
  • The Trade-off: The biggest, most powerful AI models (like the 7-billion-parameter "omniASR") were the best at understanding random, out-of-domain speech (like radio clips they hadn't seen before). However, they are so heavy they need a massive computer to run. The smaller, custom-trained models the team built are much lighter and can run on a regular laptop or phone, but they struggle a bit more with completely new types of speech.

5. The Gift: Giving Everything Away

The most important part of this paper is that the team didn't keep the keys to themselves. They opened the doors and gave away:

  • The 66 hours of recordings.
  • The transcribed text.
  • The trained AI models (the "students" who learned to speak Puno Quechua).

Why This Matters

This isn't just about making a computer understand a language; it's about dignity and access. Currently, if a Puno Quechua speaker wants to use a voice assistant, they are forced to speak Spanish, a language they might not be fully literate in. By building these tools, the researchers are handing the community a microphone that finally understands their voice, allowing them to interact with technology in their own tongue.

In short: They built the first massive library of Puno Quechua voices, taught three different AI students how to listen to them, discovered that "rough drafts" and "listening camps" are the best teachers, and then gave the whole library and the trained students to the world for free.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →