InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos
The paper introduces InteracVid, the first open-source large-scale dataset comprising over 454,000 context-query-response triplets extracted from live-chat videos, which provides the critical interaction-structured supervision needed to train multimodal models for generating natural, causal audio-visual responses to external user stimuli.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're teaching a robot to be the ultimate conversational partner. So far, we've mostly taught robots to talk using just text, like a very smart chatbot. But real life isn't just words on a screen; it's a full sensory experience. When you ask a friend a question, they don't just reply with a sentence; they might raise an eyebrow, laugh, wave their hands, or even pick up an object to show you. This paper lives in the world of multimodal AI, which is the fancy term for building computers that can see, hear, and speak all at once. The big challenge has been that while we have great tools to make robots describe things (like "a cat sitting on a mat"), we haven't had enough training data to teach them how to react to a specific person's question in real-time with the right facial expressions, gestures, and voice. It's the difference between reading a script and actually improvising a scene with a partner.
Enter InteracVid, a massive new dataset created by researchers at Tsinghua University that acts like a giant library of "real-life reactions." Instead of just describing a video, this dataset captures the exact moment a streamer (a person broadcasting live video) hears a question from a viewer and immediately responds with a smile, a gesture, or a spoken answer. The researchers scraped over 59,000 live-stream videos and extracted more than 454,000 of these perfect "question-and-reaction" moments. They found that in the wild, streamers react to everything from "What game is this?" to "How does your back hurt?" with a full-body, audio-visual response.
The team built a clever two-step system to turn this messy data into a training tool. First, they used AI to figure out which parts of a long stream were actual reactions to a chat comment, and which parts were just the streamer talking to themselves. For videos where the chat history was missing, they used AI to guess what the question might have been based on the streamer's reaction, creating a "reconstructed" question that still matched the real video response. They ended up with a dataset containing 39,000 real question-and-answer pairs and 414,000 reconstructed ones, covering everything from cooking shows to gaming sessions.
When they tested this dataset on a new generation of AI models, the results were promising but nuanced. They found that simply feeding the AI this "interaction data" made it much better at generating videos that looked and sounded like a real person responding to a user. Specifically, fine-tuning the video generator on InteracVid improved the quality of the output significantly, making the lip-syncing and gestures much more natural. However, the researchers also discovered a bottleneck: the hardest part wasn't making the video look good, but figuring out what to say or do in the first place. Even with the best video generator, if the AI didn't understand the user's question correctly, the response was off. Their experiments suggest that while we can now teach robots to perform a reaction well, teaching them to plan the right reaction is still the biggest hurdle. The paper concludes that InteracVid is a crucial first step, providing the "ground truth" of human interaction needed to build AI avatars that don't just look real, but actually feel like they're listening and responding to you.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.