← Latest papers
⚡ electrical engineering

Computational Narrative Understanding for Expressive Text-to-Speech

This paper introduces LibriQuote, a large-scale dataset of 5.3K hours of expressive character quotations from audiobooks enriched with contextual pseudo-labels, which demonstrates that fine-tuning or training on this data significantly enhances the expressivity and intelligibility of text-to-speech models.

Original authors: Gaspard Michel, Elena V. Epure, Christophe Cerisara

Published 2026-04-22
📖 5 min read🧠 Deep dive

Original authors: Gaspard Michel, Elena V. Epure, Christophe Cerisara

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are listening to a story being read aloud. Sometimes the narrator speaks in a calm, flat voice, just telling you what happened next. But then, a character speaks! Suddenly, the voice changes. It might whisper a secret, shout in anger, or giggle with joy.

For a long time, computers trying to read stories aloud (Text-to-Speech or TTS) have been great at the "calm narrator" part but often sounded like robots when trying to act out the characters. They could say the words, but they couldn't capture the feeling.

This paper introduces a new project called LibriQuote to fix that. Here is the story of what they did, explained simply:

1. The Problem: The "Robot Narrator"

Think of existing computer voices like a very polite, monotone librarian reading a book. They are clear and accurate, but if the story says, "The dragon roared!" the computer just says, "The dragon roared," in the same boring voice. It lacks the drama.

The researchers realized that human-read audiobooks are full of these dramatic moments. When a human reads a novel, they naturally switch between the "narrator" voice and the "character" voices. They wanted to teach computers to do the same thing.

2. The Solution: A New "Acting School" Dataset

The team created a massive new library of audio called LibriQuote.

  • What is it? It's a collection of over 5,000 hours of audio, but they didn't just grab random sentences. They specifically cut out the parts where characters speak (the dialogue).
  • The Secret Sauce: They didn't just take the audio; they also looked at the text right next to it. In books, authors often write clues like "he whispered softly" or "she screamed loudly."
  • The Analogy: Imagine giving a student actor a script. Most TTS systems just give them the line: "I'm scared!"
    • LibriQuote gives them the line plus the director's note: "I'm scared!" (whispered softly, trembling).
    • This extra note is called a "pseudo-label." It tells the computer exactly how to say the words.

3. The Experiments: Teaching the Computer to Act

The researchers tried two different ways to teach computers using this new dataset:

  • Method A: The "Fine-Tuning" Approach (Polishing a Pro)
    They took a smart computer voice that was already good at reading and gave it a crash course using LibriQuote.

    • Result: The computer became much clearer and easier to understand, but it didn't become a great actor yet. It was like a good student who learned the rules but hadn't found their creative spark.
  • Method B: The "From Scratch" Approach (Building a New Actor)
    They built a new computer voice from the ground up, training it only on these dramatic character lines.

    • Result: This new voice was very expressive! It could whisper and shout. However, because it had less data to learn from, it sometimes stumbled over the words (it was less clear).
  • The Breakthrough: They found that a specific type of computer model (called F5-TTS) was like a sponge. When they used LibriQuote to train it, it became both clear and expressive. It learned to act and speak clearly at the same time.

4. The Big Test: The "TTS Olympics"

To see if their new dataset actually worked, they held a competition. They took the best computer voices in the world today and asked them to read the LibriQuote test set.

  • The Verdict: Most of the top computers still sounded a bit flat or got the emotions wrong. They might say "I'm angry!" but sound happy.
  • The Winner: One system called IndexTTS2 did the best job. It was able to guess the right emotion just by looking at the context (the story around the line). However, even the winner sometimes got it wrong, proving that this is still a very hard challenge.

5. Why This Matters

This paper is like handing a new, super-detailed script to the computer voice industry.

  • For Listeners: In the future, audiobooks read by AI could sound just as dramatic and emotional as a human actor, making stories more immersive.
  • For Researchers: They now have a "gym" (the dataset) to train AI to understand the difference between a narrator's voice and a character's voice.

The Catch (Limitations)

The researchers are honest about the flaws:

  • The Source: The audio they used came from volunteers (amateur readers), not Hollywood voice actors. So, sometimes the "human" examples weren't perfect either.
  • The Bias: They don't know the gender or accent of every reader in their dataset, so the AI might accidentally learn to sound like only one type of person.
  • The Future: They warn that while AI is getting better, we still need to be careful about using it to replace real human actors, as that could hurt jobs and lead to fake voices being used for bad things (like scams).

In a nutshell: The researchers built a giant library of "character voices" with special notes on how to act. They used it to teach computers to stop sounding like robots and start sounding like storytellers. It's a huge step forward, but the computer actors still have a lot of practice to do before they can beat the best human narrators.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →