← Latest papers
💻 computer science

Talking Together: Synthesizing Co-Located 3D Conversations from Audio

This paper presents a novel dual-stream system that synthesizes realistic, spatially aware 3D facial animations for two co-located participants from mixed audio by modeling their dynamic spatial relationships and mutual gaze, supported by a large-scale dataset of over 2 million conversational pairs.

Original authors: Mengyi Shan, Shouchieh Chang, Ziqian Bai, Shichen Liu, Yinda Zhang, Luchuan Song, Rohit Pandey, Sean Fanello, Zeng Huang

Published 2026-03-10
📖 4 min read☕ Coffee break read

Original authors: Mengyi Shan, Shouchieh Chang, Ziqian Bai, Shichen Liu, Yinda Zhang, Luchuan Song, Rohit Pandey, Sean Fanello, Zeng Huang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to recreate a real-life conversation between two friends, but you only have a recording of their voices mixed together in one audio file. Most current technology is like a bad video call: it generates two separate "talking heads" floating in a void, looking straight ahead, completely ignoring each other. They might be talking, but they aren't interacting.

This paper introduces a new system called "Talking Together" that fixes this. It takes that single, mixed audio recording and generates a complete, realistic 3D scene where two people are standing face-to-face, looking at each other, nodding, and reacting naturally.

Here is how they did it, broken down into simple concepts:

1. The "Two-Stream" Kitchen

Think of the system as a kitchen with two chefs working side-by-side.

  • The Problem: Usually, if you give a chef a mixed pot of soup (the audio), they can't tell which ingredients belong to which dish.
  • The Solution: This system uses a Dual-Stream Architecture. Imagine two chefs (one for Person A, one for Person B) working in the same kitchen. They have a special "cross-talk" system. While Chef A is cooking, they can peek at what Chef B is doing to see if they need to pause, nod, or look surprised. This allows the system to figure out who is speaking and who is listening, even when both are talking at once.

2. The "Role-Playing" Badges

To help the chefs know who is doing what, the system gives them digital badges.

  • Speaker vs. Listener: The system analyzes the audio to guess who is talking and who is listening. It then gives the "Speaker" a badge that says, "You are talking, move your lips!" and the "Listener" a badge that says, "You are listening, nod and make eye contact!"
  • The Magic: Even if the audio is messy or both people talk at the same time, these badges help the system keep the two characters' animations distinct and realistic.

3. The "Eye-Contact" Coach

One of the hardest parts of a real conversation is looking at the other person. Old methods often made characters stare blankly at the camera.

  • The Fix: The researchers added a special "Eye Gaze Loss." Think of this as a strict coach standing over the animation, checking every frame. If the characters aren't looking at each other when they should be, the coach gives them a "penalty." This forces the AI to learn the subtle art of mutual eye contact and natural glances away, making the conversation feel alive.

4. The "Text-to-Scene" Director

In the past, you couldn't tell the computer where the people should stand. They just appeared in random spots.

  • The Innovation: This system lets you type a simple instruction, like "They are arguing across a table" or "They are whispering intimately."
  • How it works: The system uses a Large Language Model (like a smart assistant) to translate your text into 3D coordinates. It acts like a director telling the actors, "Okay, you stand here, and you stand there," before the animation even starts.

5. The "Data Diet"

AI models are like athletes; they need to eat a lot of data to get strong. The problem was that there wasn't enough high-quality video of two people talking in the same room.

  • The Feast: The team created a massive diet for their AI:
    1. The "Wild" Buffet: They scraped over 2 million video clips of real people talking in the wild to learn how humans actually interact.
    2. The "Sterile" Supplement: They also created a synthetic dataset by taking clean, high-quality videos of single people and mixing them together artificially. This taught the AI perfect lip-syncing (matching mouth movements to words) without the blur and noise of real-world videos.
  • The Result: By training on both, the AI learned to be both socially smart (from the wild videos) and technically precise (from the clean videos).

The Bottom Line

The result is a system that doesn't just make two people talk; it makes them converse. It captures the "dance" of a real conversation: the head tilts, the eye contact, the shared space, and the reaction to what the other person is saying. It moves us from a "video conference" feel to a "real-life meeting" feel, all generated from a single audio file.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →