← Latest papers
⚡ electrical engineering

Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens

This paper introduces SODA, a suite of native audio foundation models that jointly model semantic, acoustic, and text tokens, and establishes a validated training recipe and scaling laws through an extensive empirical study of 64 models to demonstrate superior performance across diverse audio and cross-modal tasks.

Original authors: Potsawee Manakul, Woody Haosheng Gan, Martijn Bartelds, Guangzhi Sun, William Held, Diyi Yang

Published 2026-02-19
📖 5 min read🧠 Deep dive

Original authors: Potsawee Manakul, Woody Haosheng Gan, Martijn Bartelds, Guangzhi Sun, William Held, Diyi Yang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to teach a robot to understand and speak like a human. For a long time, scientists tried two main ways to do this, but both had a major flaw:

  1. The "Translator" Approach: They took a super-smart text robot (like a chatbot) and tried to bolt a microphone onto it. The problem? The robot was so focused on the meaning of words that it forgot how to sound natural. It sounded like a robot reading a script, missing the emotion, the breath, and the unique voice of the speaker.
  2. The "Mime" Approach: They built a robot that only listened to sound patterns but ignored the actual words. It could mimic a voice perfectly but couldn't understand what was being said or hold a conversation.

This paper introduces a new kind of robot called SODA (Scaling Open Discrete Audio). Think of SODA as a "Universal Audio Chef" that learns to cook with three different ingredients at the same time: Meaning (what is said), Sound (how it sounds), and Text (the written words).

Here is how they did it, explained simply:

1. The Secret Sauce: Mixing the Ingredients

The researchers realized that to make a truly smart audio robot, you can't just feed it one type of data. You have to mix them up, sentence by sentence.

  • The Old Way: Feed the robot a book, then feed it a recording.
  • The SODA Way: Feed it a sentence of text, then immediately the sound of that sentence, then the next sentence, then the next sound. It's like teaching a child by showing them a picture of a dog, saying "dog," and then playing a bark, all in one continuous stream.

They found the perfect recipe:

  • 95% Audio: Real human speech (from YouTube and audiobooks).
  • 5% Text: High-quality written stories.
  • The Magic Ratio: They discovered that adding just a tiny bit of text (5%) helped the robot understand what people were saying without ruining its ability to understand how they sounded.

2. The "Goldilocks" Rule (Scaling Laws)

In the world of AI, there's a famous rule called "Chinchilla" for text robots: "If you double the brain size, you should double the amount of reading material."

The researchers asked: "Does this rule work for sound?"

They built 64 different versions of their audio robot, ranging from tiny to huge, and tested them with different amounts of data. They found that sound is "denser" with information than text.

  • The Analogy: Imagine text is like a clear, slow-moving river. Sound is like a fast, churning waterfall. You need a much bigger bucket (more data) to catch the same amount of water (information) from the waterfall.
  • The Discovery: To get the best audio robot, you need to increase the data much faster than you increase the brain size. Specifically, for every step up in brain size, you need about 1.6 times more data.

3. Starting from Scratch vs. Copying a Friend

Many AI researchers start by taking a pre-trained text robot and trying to teach it audio. The paper calls this "Warm-Starting."

  • The Result: It was messy. The robot got confused, had "tantrums" (training errors), and actually got worse at understanding speech.
  • The Better Way: They started with a blank slate ("Cold-Start"). They built the robot from zero, learning audio and text together from the very first second.
  • The Analogy: It's like trying to teach a professional pianist to play the drums. They might keep trying to play piano chords on the drum kit. It's better to start with a fresh student who learns drums and piano simultaneously from day one.

4. What Can SODA Do?

Because they taught it this "Universal" way, SODA is incredibly flexible. You don't need to build a new robot for every task. You just give it a new instruction, and it adapts.

  • Speech-to-Text: It can listen to a messy recording and write it down perfectly.
  • Text-to-Speech: It can read a text and speak it in a natural, human voice.
  • Voice Translation: This is the coolest trick. You can speak Spanish, and SODA can translate it to English while keeping your original voice. It's like a translator who speaks your exact tone and style, not just the words.

The Big Takeaway

This paper proves that if you treat audio, speech, and text as one big, mixed-up puzzle rather than separate subjects, you can build a much smarter, more flexible AI.

They released their "recipe," their "ingredients," and the "kitchen" (the code and models) for everyone to use. This means other scientists can now build their own audio robots without having to reinvent the wheel, potentially leading to better voice assistants, more accessible tools for the deaf or hard of hearing, and more natural human-computer interactions.

In short: They stopped trying to force a text brain to hear, and instead built a brain that was born to listen, speak, and read all at once.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →