Tagarela - A Portuguese speech dataset from podcasts
This paper introduces TAGARELA, a publicly released, large-scale Portuguese speech dataset comprising over 8,972 hours of podcast audio, designed to address resource scarcity and enable state-of-the-art automatic speech recognition and text-to-speech models for the language.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of computer voices (like Siri or Alexa) as a massive library. For languages like English, this library is a towering skyscraper filled with millions of perfect books, recordings, and scripts. Computers can learn to speak and understand English incredibly well because they have so much high-quality material to study.
But for Portuguese, the library was more like a dusty, half-empty shed. There were some books, but they were small, messy, or written in a way that computers couldn't easily read. This meant that Portuguese voices often sounded robotic, or the computers struggled to understand what people were saying.
Enter "TAGARELA."
The authors of this paper decided to build a new, massive library specifically for Portuguese. They call it TAGARELA (which sounds like "chatter" or "babble" in Portuguese). Here is the story of how they built it, explained simply:
1. The Raw Material: A Mountain of Podcasts
They started with a huge collection called "Cem Mil Podcasts" (One Hundred Thousand Podcasts). Think of this as a giant, unorganized pile of raw audio recordings.
- The Problem: These recordings were messy. Multiple people were talking over each other, there was background noise (like coffee cups clinking or traffic), and the text transcripts (the written words) were often wrong or missing. It was like trying to learn a language by listening to a crowded party where everyone is shouting at once.
- The Goal: They wanted to turn this messy pile into a pristine, organized library of 8,972 hours of clear speech. That is a lot of audio—enough to rival the biggest English datasets.
2. The Construction Crew: The "Cleaning Pipeline"
To fix the mess, the team built a digital assembly line (a pipeline) with five specialized workers:
- The Organizer (Standardization): First, they made sure every audio file looked the same, like converting all books to the same font size and paper quality.
- The Bodyguard (Speaker Separation): In podcasts, people talk over each other. This worker acts like a bouncer, separating the voices so that each recording only has one person speaking. This is crucial for teaching computers to mimic a single, clear voice.
- The Noise Filter (Overlap Detection): Even after separating speakers, sometimes two people talk at the exact same time. The team trained a smart AI to spot these "duets" and throw them away, keeping only the solo performances.
- The Translator (Transcription): They needed to turn the audio into text. They didn't just hire one translator; they used a "smart team" approach. They used a super-smart AI (Whisper) that had been trained on high-quality data to write down the words. To make sure the AI didn't "hallucinate" (make things up), they cross-checked its work with another AI. If both agreed, the text was kept.
- The Polisher (Denoising): Finally, they used a special tool to scrub away the background hiss, static, and echo, leaving behind crystal-clear audio.
3. The Result: Two Special Libraries
Once the cleaning was done, they split the collection into two special sections:
- The "Real Talk" Section (8,972 hours): This includes all the natural pauses, "umms," and "ahhs" people make when they speak. This is perfect for teaching computers to understand human speech (Automatic Speech Recognition).
- The "Studio" Section (2,800 hours): This is the super-clean, single-speaker audio. This is perfect for teaching computers to speak back to us naturally (Text-to-Speech).
4. The Test Drive
To prove their new library worked, they built two new computer models using only this data:
- The Listener: They trained a model to understand Portuguese. It became so good that it outperformed existing giants, understanding speech with very few mistakes.
- The Speaker: They trained a model to generate speech. The result was a voice that sounded very human, with a natural rhythm and flow, scoring high marks from human listeners.
Why This Matters
Before this, Portuguese speech technology was stuck in the "small shed" phase. TAGARELA is like opening the doors to a skyscraper. It gives researchers and developers a massive, high-quality foundation to build better, more natural, and more accurate voice technologies for millions of Portuguese speakers around the world.
In short: They took a messy, noisy pile of podcasts, cleaned it up with a high-tech assembly line, and turned it into the gold standard for Portuguese voice AI.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.