← Latest papers
⚡ electrical engineering

Qwen3-TTS Technical Report

The Qwen3-TTS series introduces a family of advanced, Apache 2.0-licensed multilingual text-to-speech models trained on over 5 million hours of data, featuring state-of-the-art voice cloning, fine-grained control, and a dual-track architecture with specialized tokenizers that enable real-time, ultra-low-latency streaming synthesis.

Original authors: Hangrui Hu, Xinfa Zhu, Ting He, Dake Guo, Bin Zhang, Xiong Wang, Zhifang Guo, Ziyue Jiang, Hongkun Hao, Zishan Guo, Xinyu Zhang, Pei Zhang, Baosong Yang, Jin Xu, Jingren Zhou, Junyang Lin

Published 2026-01-23
📖 5 min read🧠 Deep dive

Original authors: Hangrui Hu, Xinfa Zhu, Ting He, Dake Guo, Bin Zhang, Xiong Wang, Zhifang Guo, Ziyue Jiang, Hongkun Hao, Zishan Guo, Xinyu Zhang, Pei Zhang, Baosong Yang, Jin Xu, Jingren Zhou, Junyang Lin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a magical recording studio in your pocket that can instantly become any voice you can describe, or copy a voice from just a few seconds of audio. That's essentially what the Qwen3-TTS team is introducing in this report.

Think of this technology not just as a "text-to-speech" tool, but as a universal voice shapeshifter. Here is a breakdown of how it works and why it's special, using simple analogies.

1. The Two "Voice Engines" (Tokenizers)

The team built two different types of "engines" to turn text into sound, depending on what you need. They call these Tokenizers.

  • The "High-Fidelity" Engine (25Hz):

    • Analogy: Think of this like a high-resolution 4K movie. It captures every tiny detail of the voice, making it sound incredibly rich and natural.
    • How it works: It breaks speech into small chunks. Because it needs to look slightly ahead to ensure the sound flows perfectly (like a director checking the next scene), it takes a tiny bit more time to start the first sound.
    • Best for: When you want the absolute best quality and don't mind a split-second delay.
  • The "Instant" Engine (12Hz):

    • Analogy: Think of this like a high-speed courier service. It doesn't wait to see the whole package before sending the first box; it sends the first box the moment it's ready.
    • How it works: It uses a clever trick where it sends the most important "meaning" of the words first, then fills in the acoustic details immediately after. This allows it to start speaking in just 97 milliseconds (faster than a human blink).
    • Best for: Real-time conversations, like talking to a robot assistant where you don't want to wait for it to "think" before it speaks.

2. The "Voice Cloning" Magic

Usually, cloning a voice takes hours of recording. Qwen3-TTS changes the game by saying, "Give me 3 seconds, and I'll give you the whole voice."

  • The Analogy: Imagine you have a master chef who can taste a single spoonful of soup and instantly recreate the entire recipe, including the secret spices.
  • What it does: You can feed it a 3-second clip of someone talking, and it can generate unlimited new sentences in that exact voice. It can also create brand new voices just by you describing them (e.g., "a deep, grumpy robot voice that sounds like it just woke up").

3. The "Multilingual Polyglot"

This model isn't just good at one language; it's a linguistic chameleon.

  • The Analogy: Imagine an actor who has memorized a script in 10 different languages but keeps the same personality and voice acting style in all of them.
  • What it does: It was trained on over 5 million hours of speech data across 10 languages. If you ask it to speak Korean using a voice it learned from a Chinese speaker, it doesn't just translate the words; it keeps the original speaker's unique "timbre" (the texture of their voice) while speaking Korean perfectly. It handles tricky language pairs (like Chinese to Korean) better than any previous system.

4. The "Infinite Storyteller"

One of the biggest problems with AI voices is that they often get tired, repeat themselves, or sound robotic after a few minutes.

  • The Analogy: Think of a tired radio host who starts stumbling over words after an hour. Qwen3-TTS is like a super-energized narrator who can read a 10-minute story without losing their breath or changing their tone.
  • What it does: The team trained it specifically to handle long texts. It can generate over 10 minutes of continuous, fluent speech without the glitches or "glitches" that usually happen when AI tries to speak for too long.

5. The "Instruction-Following" Director

You aren't just limited to copying voices; you can direct the performance.

  • The Analogy: Imagine you are a movie director talking to an actor. You can say, "Read this line, but make it sound like you're whispering a secret," or "Say this, but sound like you're excited about a surprise."
  • What it does: You can type instructions like "speak slowly and sadly" or "sound like a cheerful news anchor," and the model adjusts the voice instantly to match your description. It understands these complex instructions better than many previous models.

The Bottom Line

The Qwen3-TTS team has released a family of models that are fast, multilingual, and incredibly controllable. They have solved the "speed vs. quality" trade-off by offering two versions (one for instant speed, one for maximum quality) and have proven that AI can now clone voices from tiny samples, speak fluently in multiple languages simultaneously, and tell long stories without getting tired.

They have made these tools available to everyone (open source), hoping that developers and researchers can use them to build the next generation of human-computer conversations.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →