ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
The paper introduces ParsVoice, a publicly available 2,200-hour multi-speaker Persian speech corpus constructed via a scalable pipeline from audiobooks, which significantly expands open Persian TTS resources and demonstrates high-quality synthesis performance when used to fine-tune the XTTS model.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you want to teach a robot to speak Persian. You can't just give it a dictionary; you need to feed it thousands of hours of real human voices reading stories, so it can learn the rhythm, the emotion, and the unique sounds of the language.
For a long time, the Persian language has been like a guest at a massive party who wasn't given a seat at the table. While languages like English have huge libraries of recorded speech (think of them as massive, well-organized warehouses), Persian had very little. This made it hard to build good voice assistants or storytelling robots for Persian speakers.
The Solution: ParsVoice
The authors of this paper built ParsVoice, which is essentially a massive, high-quality library of Persian speech. It's the biggest public collection of its kind ever made for this language.
Here is how they did it, broken down into simple steps:
1. The Raw Material: Audiobooks
Instead of hiring actors to read scripts in a studio (which is expensive and slow), the team went to a public library of audiobooks called IranSeda. Think of this as finding a mountain of raw, uncut film reels. They had over 3,800 books, but there was a catch: they didn't have the written scripts (transcripts) that matched the audio perfectly.
2. The "Smart Cutter" Pipeline
You can't just take a 10-hour audiobook file and feed it to a robot; it needs to be chopped into tiny, perfect sentences. The team built an automated "factory line" to do this:
- The Rough Cut: They used a computer program to find the silence between sentences and cut the audio there.
- The "Did You Finish?" Check: Sometimes, the computer cuts the audio in the middle of a sentence (like stopping a movie right before the hero says "I love you"). They used a smart AI (based on a model called ParsBERT) to read the text and ask, "Is this a complete thought?" If the answer was "No," the system automatically extended the cut by a fraction of a second and checked again until the sentence was whole.
- The "Trim the Fat" Step: Even after cutting, there might be a tiny bit of silence or a cough at the start or end of a clip. The system used a "binary search" (a clever way of guessing and checking) to snip off exactly the right amount of silence so the voice starts and stops cleanly, without chopping off any words.
- The Quality Inspector: Not all audiobooks are recorded in a soundproof studio. Some might have background music or bad microphones. The system acted like a strict editor, scoring every clip on audio clarity and text accuracy. If a clip was too noisy or the text was gibberish, it was thrown out.
- The Voice Detective: Since one audiobook might have multiple narrators, the system used a "voice fingerprint" tool to group all the clips by who was speaking. This allowed them to identify 1,815 different speakers automatically.
3. The Result
The final product is a 2,200-hour library of clean, perfectly chopped Persian sentences.
- It is 25 times larger than the previous biggest Persian speech collection.
- It contains 1.36 million individual sentence clips.
- It covers 1,815 different voices.
4. Did It Work?
To test if this library was actually useful, the team taught a modern AI voice model (called XTTS) using only this new data. They didn't teach it the sounds of Persian letters (phonemes) first; they let it learn directly from the text and audio.
The results were impressive:
- Naturalness: Humans rated the robot's voice as sounding very natural (3.6 out of 5).
- Voice Matching: When they asked the robot to mimic a new voice it had never heard before, it did a great job (4.0 out of 5).
- Intelligibility: The robot spoke clearly enough that other computer programs could understand it perfectly.
What It Doesn't Do (The Limitations)
The authors are honest about what this library isn't:
- It's not for chatting: Because the data comes from audiobooks, the voices sound like they are reading a story (formal and steady). If you ask the robot to have a casual, fast-paced chat like two friends at a coffee shop, it might sound a bit stiff.
- It's not perfect: Since they had to use a computer to guess the text (because they didn't have the original books), there are a few tiny mistakes in the text, though they filtered out most of them.
- Gender Balance: The library has more male voices than female voices, simply because the audiobooks they used had more male narrators.
In Summary:
The team built a giant, automated factory that turned thousands of hours of raw Persian audiobooks into a clean, organized, and massive dataset. This dataset is now available for anyone to use to build better Persian voice technologies, finally giving the language a seat at the table of high-tech speech research.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.