← Latest papers
💻 computer science

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

The paper introduces ScenA, a reference-driven multi-speaker audio generation method that leverages a text-to-audio flow-matching model pretrained on in-the-wild data to create realistic, overlapping dialogue scenes with ambient noise by conditioning on reference voices and free-form prompts, while overcoming the "Reference Shortcut" challenge through a high-noise-biased training strategy.

Original authors: Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen

Published 2026-06-19
📖 4 min read☕ Coffee break read

Original authors: Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to create a short, lively radio play with two different actors, background music, and the sound of a busy café.

The Old Way (The "Assembly Line" Approach)
Previously, if you wanted to make this scene, you had to act like a strict film director with a clipboard. You would have to:

  1. Write a script that says exactly who speaks when (e.g., "Line 1: Speaker A," "Line 2: Speaker B").
  2. Generate the voice for Speaker A separately.
  3. Generate the voice for Speaker B separately.
  4. Manually glue the clips together.
  5. Add the background noise on top.

The result often sounded "clean" but fake, like two people talking in a vacuum, because the system didn't know how they should overlap, laugh together, or react to the room's acoustics.

The New Way (SCENA: The "Improvisational Director")
The paper introduces a new system called SCENA. Instead of a clipboard, you just give the AI a free-flowing story description in plain English.

  • The Prompt: You type: "Soft piano plays. The first speaker says 'Hi!' quickly. The second speaker shouts 'Welcome!' at the same time, and their voices mash together perfectly."
  • The References: You upload two short audio clips of the voices you want to use (Voice 1 and Voice 2).
  • The Magic: The AI reads your story and instantly generates the whole scene. It figures out who speaks when, makes the voices overlap naturally, adds the piano, and even makes the room sound real—all in one go.

The Big Problem They Solved: "The Cheating Shortcut"

When the researchers first tried to teach the AI this way, it tried to "cheat."

Imagine a student taking a test where they have to match a photo of a person to a description.

  • The Goal: The student should read the description ("The man with the red hat") and find the matching photo.
  • The Cheat: The student looks at the photo they are trying to solve, sees the red hat, and just picks the photo that looks most like it, ignoring the description entirely.

In the AI's case, during training, the "test" was a noisy, scrambled version of the audio. The AI realized it could just listen to the scrambled noise, find the voice that sounded most similar, and copy it. It didn't need to read your text prompt to know who was speaking. This worked great during training (low error scores) but failed miserably when you actually used it, because real generation starts with pure silence (no noise to cheat with). The AI had learned to ignore your instructions.

The Solution: "The High-Noise Diet"

To fix this, the researchers changed the "training diet" for the AI.

Normally, AI training happens at a "medium noise" level where the answer is still somewhat visible. The researchers realized this was where the cheating happened. They switched to a High-Noise-Biased schedule.

Think of it like teaching someone to recognize a face in a blizzard:

  1. Old Method: Show them a slightly foggy photo. They can still see the features and guess the name without reading the clue.
  2. New Method (SCENA): Show them a photo that is 90% white snow. The only way to guess who it is is to read the clue ("It's the guy with the red hat").

By forcing the AI to learn mostly when the audio is almost pure static, they forced it to stop cheating and actually learn to listen to your text instructions to decide which voice speaks when.

The Results

When they tested this new system:

  • Better Binding: It correctly assigned the right voice to the right part of the story much better than previous systems.
  • Realism: It created overlapping speech, laughter, sighs, and background noise that sounded like a real, messy conversation, not a sterile studio recording.
  • Wild Cards: It even worked well when the voice clips were recorded in noisy, real-world environments (like a street or a park), not just in a quiet studio.

In short, SCENA is a system that lets you direct a multi-actor audio scene using just a story and some voice clips, without needing to write a complex script or manually edit the audio, because it learned to listen to your words instead of looking for shortcuts.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →