← Latest papers
💻 computer science

Conversational Human Audio-visual Talking Dialogue Generation

The paper introduces CHAT, a novel framework that generates diverse, aligned dyadic audio-visual dialogue clips from single text prompts by unifying large language and talking face models, thereby offering a scalable and ethically sound alternative to costly real-world data collection while demonstrating superior performance and utility for downstream tasks.

Original authors: Junhao Song, Lluis Guasch, Xilin He, Zhongyu Yang, Yingfang Yuan, Weicheng Xie, Linlin Shen, Haijun Lin, Shizhe Liu, Wei Pang, Siyang Song

Published 2026-07-07
📖 4 min read☕ Coffee break read

Original authors: Junhao Song, Lluis Guasch, Xilin He, Zhongyu Yang, Yingfang Yuan, Weicheng Xie, Linlin Shen, Haijun Lin, Shizhe Liu, Wei Pang, Siyang Song

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to create a movie scene where two people are having a lively conversation. Usually, to get this right, you need to hire actors, set up cameras, record audio, and spend hours editing to make sure their lips move with the words and their facial expressions match the mood. It's expensive, time-consuming, and sometimes raises privacy concerns.

This paper introduces a new AI tool called CHAT that acts like a "magic scriptwriter and director" rolled into one. Instead of hiring actors, you simply type a single sentence describing a scenario (like "Two friends discussing AI"), and CHAT creates the entire video clip for you.

Here is how it works, broken down into simple steps:

1. The Scriptwriter (Textual Dialogue Generation)

First, the AI acts as a creative writer. You give it a prompt, and it invents a full conversation between two characters. It doesn't just write the words; it also invents who these people are (e.g., "a thoughtful gentleman" and "a vibrant lady") and decides how they should feel (happy, serious, excited).

2. The Voice Actor & Director (Audio Refinement)

Next, the AI turns that text into speech. But it doesn't just read the words robotically. It adds a layer of "human touch":

  • Emotion: It makes sure the voices sound happy, sad, or angry depending on the context.
  • Realism: In real life, when one person talks, the other often makes small noises like "Hmm," "Uh-huh," or "Oh!" to show they are listening. This AI adds those little sounds automatically, so the conversation feels like a real back-and-forth, not two people taking turns reading a script.

3. The Animator (Facial Behaviour Generation)

Finally, the AI brings the characters to life visually. This is the most complex part, and it uses a clever "two-step dance":

  • Step A (The Basics): It first makes the characters' mouths move to match the words they are saying.
  • Step B (The Reaction): This is where the magic happens. The AI looks at what the other person is doing and saying. If Person A is telling a joke, the AI makes Person B smile or laugh while Person A is talking. If Person A is sad, Person B looks concerned. It ensures that the two faces are reacting to each other in real-time, creating a "mutual responsiveness" that most other AI tools miss.

Why is this a big deal?

The authors explain that existing AI tools usually do one of two things:

  1. Talk alone: They make a person talk, but they don't react to anyone else.
  2. React alone: They make a person react to a video, but they don't generate the conversation or the other person's speech.

CHAT is unique because it generates both sides of the conversation at the same time, with both people reacting to each other naturally.

The "50,000 Scene" Library

To prove this works, the researchers used CHAT to automatically generate a massive library of 50,000 different conversation clips (called the CHAT-AVD-50k dataset). They say this is like creating a giant training gym for other AI systems. When they used this library to train other AI models, those models got better at understanding human interactions than when they were trained only on real human videos.

The Catch

The paper admits that this process is currently slow and expensive to run. It requires a lot of computing power because it's essentially running a writer, a voice actor, and an animator simultaneously for every single clip. It's not yet a tool you can run on a laptop in seconds, but it proves that creating realistic, interactive human conversations from scratch is possible.

In short: CHAT is a system that can take a simple idea and turn it into a realistic video of two people having a natural, emotional, and interactive conversation, solving the problem of how to get "two actors" to perform perfectly together without ever needing to hire them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →