UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception
UniAudio-Token enhances semantic speech tokenizers with general audio perception capabilities by introducing Semantic-Acoustic Primitives for structured supervision and a Semantic-Acoustic Equilibrium gating mechanism to restore acoustic details, thereby achieving a unified audio interface that outperforms existing single-codebook baselines in both understanding and generation tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a super-smart robot (an AI) how to understand and speak the world of sound. To do this, the robot needs a "dictionary" to translate sound waves into words it can read. This dictionary is called a tokenizer.
For a long time, these dictionaries had a major problem: they were like specialized translators who only spoke "Human Speech." They were great at understanding what a person said, but if you played them a recording of a barking dog, a crashing wave, or a siren, they would go deaf. They would try to force those sounds into speech patterns, often getting it wrong or ignoring the details entirely. This is what the paper calls "acoustic blindness."
On the other hand, there were other dictionaries designed to hear everything (like a high-fidelity microphone), but they were terrible at understanding the meaning behind the words. They could hear the exact pitch of a voice but couldn't tell you if the person was happy or sad, or what they were actually saying.
UniAudio-Token is the new, all-in-one dictionary that fixes this. The authors (from Peking University and Tencent) built a system that acts like a universal translator for all sounds, not just human speech.
Here is how they did it, using two main "superpowers":
1. The "Three-Layer Sandwich" (Semantic-Acoustic Primitives)
Imagine you are describing a movie scene to a friend.
- Old way: You just say, "A man is walking." (This is the linguistic part). You ignore the fact that he is walking in the rain, wearing a red coat, and looks sad.
- UniAudio-Token's way: They invented a new way to describe sound called Semantic-Acoustic Primitives (SAP). It's like a three-layer sandwich:
- Layer 1 (The Script): What is being said? (The words).
- Layer 2 (The Actor): How is it being said? (Is the voice deep? Is the person angry? Is it a child or an elder?).
- Layer 3 (The Stage): What is happening around them? (Is there rain? Is there a car engine humming? Is there a door slamming?).
By forcing the AI to learn all three layers at once, it stops ignoring the "noise" (like rain or music) and starts treating it as important information.
2. The "Smart Mixer" (Semantic-Acoustic Equilibrium)
Even with a great description, there was a technical problem. As the AI processes sound deeper into its brain, it tends to throw away the "fine details" (like the texture of a voice or the echo of a room) to focus on the main meaning. It's like summarizing a book so quickly that you lose the beautiful descriptions of the scenery.
To fix this, they built a Smart Mixer (called SAE).
- Think of the AI as a factory assembly line. The early stations see the raw, detailed sound (the "shallow" layers). The later stations see the big picture meaning (the "deep" layers).
- Usually, the deep layers forget what the early layers saw.
- The Smart Mixer is a gatekeeper that watches the content. If the AI is listening to a complex song or a noisy street, the gate opens wide to let the detailed "raw sound" back in to mix with the "big picture." If the AI is listening to clear speech, the gate closes a bit to focus on the words.
- This happens automatically and instantly, ensuring the AI never loses the "flavor" of the sound while still understanding the meaning.
The Results: A Swiss Army Knife
The paper shows that this new system is a "Swiss Army Knife" of audio:
- It hears everything: When tested on sounds like dogs barking, rain falling, or helicopters flying, it groups them perfectly. Old systems would mix these up or ignore them.
- It speaks perfectly: Despite listening to all these non-speech sounds, it didn't get worse at understanding human speech. In fact, it got better at generating human speech because it learned to keep the tiny details (like accents and breath) that make speech sound real.
- It helps the big brain: When they plugged this tokenizer into a large Language Model (the "brain"), that brain became much smarter at answering questions about music, sound effects, and speech, outperforming all previous single-dictionary systems.
In short: UniAudio-Token is a new tool that teaches AI to listen to the whole world of sound—words, voices, and background noise alike—without losing the ability to understand or speak clearly. It bridges the gap between "hearing" and "understanding."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.