An Artificial Glottal Spectrum-Based Excitation Model with LP-Based Spectral Modelling for Prosody Modification
This paper proposes a novel prosody modification method that replaces traditional time-domain pitch marking with an artificial glottal spectrum-based excitation model to eliminate pitch-marking ambiguities and artifacts, thereby achieving high-quality, natural, and efficient pitch and duration manipulation suitable for text-to-speech and voice conversion applications.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Human speech is a complex physical event, a fleeting vibration of air that carries meaning, emotion, and identity. At its core, the sound we produce relies on two distinct parts working in tandem. The first is the source: the raw, rhythmic pulse created when air rushes through the vocal cords in the throat. The second is the system: the shape of the mouth, tongue, and throat that acts like a filter, coloring that raw pulse into the specific sounds of vowels and consonants. When we speak, these two elements are inextricably linked, making it difficult to change one without affecting the other. This presents a significant challenge for engineers trying to build machines that can speak naturally or alter the voice of a recording. If a computer tries to simply stretch or squeeze a recording to change the speed or the pitch of a voice, the result often sounds robotic or distorted, losing the natural quality of the human speaker.
Researchers have long sought ways to untangle these two components so they can be manipulated independently. The most common approach for decades has been a technique that cuts the speech signal into tiny pieces based on the timing of the vocal cord vibrations, then rearranges those pieces to change the speed or pitch. While effective for small adjustments, this method struggles when asked to make drastic changes. It relies heavily on finding the exact moment the vocal cords close, a task that is difficult to do perfectly. When the computer guesses wrong about these moments, the reconstructed voice develops strange artifacts, sounding metallic or glitchy, especially when the pitch is raised significantly or lowered too much.
In a recent study, a team of researchers from India proposed a different way to handle this problem. Instead of trying to salvage and rearrange the original sound waves, they decided to replace the source entirely. Their method involves stripping away the original vocal cord vibration from a speech recording and replacing it with a perfectly clean, artificial version generated by a computer. They keep the original "filter"—the unique shape of the speaker's mouth and throat—intact, ensuring the voice still sounds like the same person. By generating a new, idealized pulse based on the desired pitch and combining it with the original filter, they can reshape the voice without the distortions that plague older methods.
The researchers tested this new approach against the traditional method using recordings of Tamil speakers. They asked human listeners to rate the quality of the modified speech on a scale from one to five. The results showed that their new method consistently produced higher-quality voices, particularly when the pitch was changed by large amounts. While the traditional method began to sound unnatural and degraded when the pitch was altered beyond a certain point, the new approach maintained clarity and naturalness across a wider range. The team found that this technique worked well for both male and female voices, preserving the speaker's identity even when the pitch was shifted significantly up or down.
Beyond simple pitch changes, the study also explored how well the method could handle expressive speech, where the tone of voice rises and falls to convey emotion or emphasis. The researchers created specific patterns of pitch movement, such as rising tones for questions or falling tones for statements, and applied them to the speech. The new method handled these dynamic changes more smoothly than the traditional approach, which often introduced audible glitches when the pitch contour changed rapidly. The artificial source they generated was flexible enough to follow these complex curves without breaking the flow of the sound.
The process of rebuilding the voice from these separate parts required a specific step to ensure the sound was not just a collection of frequencies but a coherent wave. Since the researchers worked with the magnitude of the sound waves but not their timing phase, they used an iterative mathematical process to estimate the missing timing information. This step, known as phase reconstruction, allowed them to synthesize a final waveform that sounded natural to the human ear. While the researchers noted that this reconstruction step is not perfect and leaves a tiny amount of room for improvement, the overall quality of the speech was superior to existing techniques.
The study concludes that by abandoning the attempt to perfectly track the original vocal cord vibrations and instead generating a clean, artificial source, it is possible to achieve more robust and expressive voice modification. This approach removes the ambiguity and errors associated with detecting the exact moments of vocal cord closure, which has been a persistent bottleneck in the field. The findings suggest that for applications like text-to-speech systems or voice conversion, where naturalness and flexibility are paramount, replacing the source with a controlled, artificial model offers a more reliable path forward than trying to stretch the original signal. The work demonstrates that sometimes, to make a voice sound more human, one must be willing to replace the human element with a precise, mathematical ideal.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.