Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
This paper proposes a two-stage prompt selection strategy that evaluates candidates based on prosodic features, audio quality, and text-emotion coherence before synthesis, and aligns them with input text during synthesis, to significantly improve speaker consistency and emotional intensity in zero-shot text-to-speech generation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a director trying to film a movie scene where an actor needs to deliver a line with perfect sadness or explosive anger. You have a script (the text), but you need a "reference clip" (the prompt) to show the actor how to say it.
If you hand the actor a reference clip of them whispering calmly, they might whisper the sad line. If you hand them a clip of them screaming in joy, they might shout the sad line. The reference clip dictates the vibe, the voice, and the intensity of the performance.
This paper, titled "Expressive Prompting," tackles a major problem in AI voice generation: How do we pick the perfect reference clip to make the AI sound exactly how we want?
Here is the breakdown of their solution, ExpPro, using simple analogies.
The Problem: The "Random Guess" Game
Currently, most AI voice systems pick a reference clip at random or just by matching the words.
- The Random Approach: Like asking a friend to "pick a random song" to set the mood for a party. You might get a sad ballad when you wanted a dance track.
- The Text-Only Approach: Like looking at a song title ("Happy Birthday") and assuming the song sounds happy. But what if the singer is singing it in a funeral dirge? The words match, but the emotion is wrong.
The result? The AI sounds flat, the emotion is weak, or the voice sounds like a different person than the one you wanted.
The Solution: The "Two-Stage Casting Call" (ExpPro)
The authors propose a smart, two-step strategy to find the best reference clip without needing to retrain the AI (which is like teaching the actor a new skill from scratch). They call this ExpPro.
Stage 1: The "Audition" (Static Selection)
Before the AI even starts speaking, the system acts like a strict casting director holding a massive pile of reference clips. It filters them through three specific lenses:
The "Pitch" Check (The Volume Knob):
- The Analogy: Think of pitch as the "energy level" of a voice. A low, steady hum is usually calm or sad. A high, wiggly voice is usually happy or surprised.
- What they do: The system measures the average height (mean) and the wiggles (variance) of the voice in the clips. It groups them into clusters. If you want "Angry," it picks clips with high, jagged energy. If you want "Sad," it picks low, flat ones.
The "Vibe Check" (Perception & Text Match):
- The Analogy: Imagine a music critic listening to the clip. Is the audio clear and natural? Also, does the text in the clip actually match the emotion?
- What they do: They use an AI "critic" (ChatGPT) to read the text of the reference clip and ask, "Does this sentence actually sound like 'Angry'?" If the clip says "I am so happy" but sounds bored, it gets rejected.
The "Rehearsal" (Model Performance):
- The Analogy: This is the most crucial step. Just because a clip sounds good on its own doesn't mean the AI can copy it well. It's like a great singer who sounds amazing live but terrible when recorded on a cheap phone.
- What they do: The system actually tests the top clips by asking the AI to speak a few neutral sentences using that clip as a guide. It checks:
- Did the AI understand the words? (Character Error Rate)
- Did it keep the same voice? (Speaker Similarity)
- Did it keep the same emotion? (Emotion Similarity)
- Only the clips that the AI can actually copy well are kept.
Stage 2: The "Final Match" (Dynamic Selection)
After the "Audition," you have a shortlist of the best 5 or 10 clips. Now, you need to pick the one that fits the specific sentence you want the AI to say right now.
- The Analogy: You have a shortlist of 5 great actors. You are about to film a scene where the character says, "I can't believe you did that!"
- If the scene is a comedy, you pick the actor who is good at sarcastic delivery.
- If the scene is a tragedy, you pick the actor who is good at heartbreak.
- What they do: The system compares the meaning of the new sentence with the meaning of the shortlisted clips. It picks the one that is semantically closest. If the new text is about "loss," it picks the "sad" clip from the shortlist, not the "happy" one, even if both passed the first stage.
Why This Matters
The paper proves that by using this two-step "Casting Call," the AI becomes:
- More Emotional: The voices actually feel angry, sad, or happy, not just robotic.
- More Consistent: The voice sounds like the same person throughout the speech.
- Better Quality: The speech sounds clearer and more natural.
The Bottom Line
Think of ExpPro as a super-smart assistant that doesn't just grab a random voice sample. Instead, it:
- Auditions thousands of clips to find the ones with the right energy and clarity.
- Tests them to see which ones the AI can actually copy well.
- Matches the final winner to the specific mood of your sentence.
The result? An AI that can tell a story with the same emotional depth and voice consistency as a human actor, simply by choosing the right "reference" to start with.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.