Unified Audio Intelligence Without Regressing on Text Intelligence
The paper introduces Audex, a unified 30B audio-text LLM built on a strong text-only MoE backbone that achieves state-of-the-art performance across diverse audio and speech tasks while preserving the original model's advanced reasoning and agentic capabilities without significant regression.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: A "Swiss Army Knife" Brain That Doesn't Forget How to Think
Imagine you have a brilliant professor who is a master at writing essays, solving complex math problems, and coding software. This professor is your text-only AI. Now, imagine you want to teach this professor to also understand music, recognize bird songs, and speak back to you in a human voice.
Usually, when you teach a genius a new, difficult skill (like juggling), they might get a little clumsy at their original job (like writing). They might start making grammar mistakes or forgetting facts because their brain is so busy learning the new trick.
This paper introduces "Audex," a new AI model that learns to juggle audio and speech without dropping its ability to think clearly.
The researchers built Audex on top of a very smart text-only brain called Nemotron-Cascade-2. Their goal was simple: Make an AI that can hear, speak, and understand sound, but never lose its ability to reason, solve math problems, or follow instructions.
How It Works: The "Universal Translator" Trick
Most AI models that handle sound are like two separate people holding hands: one person listens, and another person speaks. If they don't talk to each other perfectly, the result is messy.
Audex is different. It is a single, unified brain. Here is how it manages to do everything at once:
- The Ear (Audio Encoder): When Audex hears a sound (like a dog barking or someone speaking), it doesn't try to understand the raw sound waves immediately. Instead, it uses a specialized "translator" (an audio encoder) to turn the sound into a list of numbers that look just like the words the AI already knows.
- The Brain (The LLM): These numbers are fed into the main brain. Because the brain sees them as "tokens" (like words), it treats a sound clip exactly the same way it treats a sentence of text.
- The Mouth (Audio Codec): When Audex wants to speak or make a sound, it doesn't just "think" the words; it predicts the next "sound token" in the sequence, just like it predicts the next word in a sentence.
The Analogy: Imagine a chef who usually only cooks with text recipes. Audex is like teaching that chef to cook with sound ingredients. Instead of needing a separate kitchen for sound, Audex just translates the sound ingredients into the same language the chef already speaks. The chef can now cook a "sound soup" without forgetting how to bake a "text cake."
The Secret Sauce: Training Without Breaking the Brain
The researchers faced a big challenge: If you train a smart text AI too hard on audio, it might forget how to be smart. To prevent this, they used a careful, step-by-step training recipe:
- Warm-up: First, they taught the AI how to handle the "sound ingredients" (the audio tokens) without changing its core knowledge.
- Layered Learning: They didn't throw everything at the AI at once. They taught it to generate sound first, then taught it to understand sound, and finally mixed it all together. This is like learning to ride a bike with training wheels before trying to ride a unicycle.
- The "No-Regression" Promise: After all the training, they tested the AI on its old skills (math, logic, coding). The results were shocking: Audex got just as good at these tasks as the original text-only version. It didn't lose any of its "genius" while gaining the ability to hear and speak.
What Can Audex Actually Do?
According to the paper, Audex is a "Swiss Army Knife" for audio. It can:
- Listen and Understand: It can answer questions about what it hears (e.g., "Is that a dog or a wolf?" or "What is the mood of this music?").
- Transcribe and Translate: It can turn spoken words into text and translate them into other languages, even in noisy rooms.
- Speak and Sing: It can take a text prompt and generate human-like speech or even create music and sound effects (like "rain falling" or "birds chirping").
- Talk to Talk: It can take a voice input and generate a different voice output (speech-to-speech).
The Results: The Best of Both Worlds
The paper compares Audex to other top-tier AI models.
- On Text: It performs as well as the smartest text-only models, solving complex math and logic puzzles without a drop in performance.
- On Audio: It beats many other models at recognizing speech and understanding general sounds.
- The "Thinking" Mode: Unlike some models that get "dumber" when they try to think and speak at the same time, Audex can "think" (reason) about a sound problem and then generate the answer in audio, all in one go.
In Summary
This paper presents Audex, a model that proves you don't have to choose between being a "smart thinker" and a "good listener/speaker." By treating sound as just another type of language, the researchers built a single AI that can hear, speak, and reason with equal brilliance, keeping its original intelligence intact while mastering the world of sound.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.