PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
The paper proposes PASE, a generative speech enhancement framework that leverages the robust phonological priors of a pre-trained WavLM model to effectively mitigate both linguistic and acoustic hallucinations while achieving superior perceptual quality compared to existing methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to listen to a friend's voice in the middle of a chaotic, roaring stadium. Your brain is a miracle of engineering; it filters out the screaming fans and the blaring music to focus on your friend's words. For decades, scientists have tried to build computers that can do the same thing, a field called Speech Enhancement. The goal is simple: take a messy, noisy recording and clean it up so it sounds like it was recorded in a quiet room.
Traditionally, computers did this by acting like a very strict editor, cutting out the "bad" parts of the sound wave. But this often made the voice sound robotic or flat. Recently, a new generation of "generative" computers has arrived. Think of these not as editors, but as talented artists who can re-imagine what the voice should sound like. They are great at making the voice sound natural and smooth. However, there is a catch: because they are so good at imagining, they sometimes get too creative. If the noise is too loud, they might invent words that were never spoken or change the speaker's voice to sound like a different person entirely. This is called "hallucination." It's like a musician trying to fix a broken recording but accidentally playing the wrong notes because they guessed what should have been there. The big question for scientists is: How do we get the natural sound of the new artists without letting them make up the story?
This is where a new study by researchers from Nanjing University and Cisco Systems comes in. They propose a system called PASE (Phonologically Anchored Speech Enhancer) to solve this "too creative" problem. Instead of teaching the computer to guess the words from scratch, they decided to give it a reference based on how human speech actually works.
The researchers realized that previous "generative" systems were trying to learn the rules of language by looking at the noisy, broken recordings. It's like trying to learn the rules of soccer by watching a game played in a blizzard where you can barely see the ball. The computer gets confused and starts inventing rules. PASE takes a different approach. It uses a pre-trained AI model called WavLM, which has already studied thousands of hours of clear speech. Think of WavLM as a master linguist who knows exactly how human sounds (phonemes) fit together. PASE doesn't try to re-learn these rules; it simply asks the master linguist for help. It uses the linguist's knowledge to "anchor" the cleaning process, ensuring the computer never invents a word that doesn't fit the sound structure of human speech.
To handle the voice's personality (like the speaker's unique tone and pitch), PASE uses a clever two-track system. Imagine the computer is a painter. One track focuses on the story (the words and meaning), using the master linguist's help to make sure the story is accurate. The other track focuses on the texture (the speaker's voice), pulling fine details from the very bottom layers of the AI to keep the voice sounding like the original person. By keeping these two tracks separate but working together, PASE avoids the common trap of changing the speaker's identity while fixing the words.
The results of their experiments are quite promising. When they tested PASE against other top models, it didn't just sound good; it was much more accurate. In tests with very noisy audio, other models often made up words, resulting in a "Word Error Rate" (a measure of how many words were wrong) as high as 36% or even 72% in extreme cases. PASE, however, kept the error rate much lower, around 7.5% to 15% depending on the test, while still sounding natural. It also did a better job of keeping the speaker's voice sounding like the original person, rather than morphing into a stranger.
The authors suggest that the key to this success wasn't just having a bigger computer or more data, but using the right "prior knowledge." They found that the ability to understand speech structure comes from the specific way the AI was trained to predict missing sounds, not just from seeing a massive amount of data. By leveraging this specific knowledge, PASE acts like a disciplined artist: it has the creativity to fill in the gaps of a noisy recording, but it is strictly guided by the rules of human language so it never crosses the line into making things up. This makes it a more reliable tool for real-world situations where clarity and accuracy are just as important as sounding natural.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.