TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics
This paper introduces TIDES, a high-resolution longitudinal bilingual dataset of 12 university teams over a semester, which demonstrates that while fine-tuning on its socio-structural annotations significantly improves next-speaker prediction, it does not necessarily enhance the naturalness or coherence of generated multi-party conversations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are walking into a bustling coffee shop. You hear a group of friends laughing, arguing, and planning a road trip. You don't need to know their names to guess who will speak next; you just listen to the rhythm. One person finishes a sentence, pauses, and another jumps in with a joke. This rhythm is called turn-taking, and it's the invisible glue that holds any conversation together. For a long time, computer programs (called Large Language Models, or LLMs) have been great at chatting one-on-one, like a text message between two people. But when you throw a whole group into the mix, these computers get confused. They struggle to understand the messy, shifting dance of a real group, where roles change, inside jokes develop, and the conversation flows differently every time. Scientists want to build AI that can join these groups naturally, but to do that, they need to study how real humans actually behave over time, not just in short, fake experiments.
This is where a new study called TIDES comes in. The researchers, a team from KAIST, decided to stop guessing and start watching. They tracked 12 real university student teams over an entire semester as they worked on their projects. Instead of giving them a fake script to follow, they just let the students do their thing, recording over 75,000 spoken sentences from 88 different meetings. They didn't just listen to what was said; they also tracked who was doing what, noting when a student became the "leader," the "critic," or the "problem solver" as the weeks went by. It's like filming a whole season of a reality show, but instead of drama, it's about how teams actually solve problems together.
The team then taught a computer model to watch these recordings and predict who would speak next. The results were surprisingly good. When the model learned from this real-world data, it got much better at guessing the next speaker—jumping from a 50% guess to a 64.5% success rate. It even beat some of the most expensive, powerful AI models available today, doing so with less than half the amount of training data. The model learned that in the early days of a team (when everyone is awkward and figuring things out), the conversation is chaotic and hard to predict. But as the team matures and finds its groove, the pattern becomes clear, and the AI can see the next speaker coming from a mile away.
However, there is a twist in the story. The researchers asked a crucial question: Just because the AI can guess who speaks next, does that mean it can write what that person says in a way that sounds natural? They tested this by having the AI generate new sentences for the conversation. Here, the results were a bit disappointing. Even though the AI was great at predicting the structure of the conversation, the sentences it actually wrote felt stiff and robotic to human listeners. People preferred the "vanilla" (untouched) AI models, which sounded more natural, even if they were worse at guessing who would talk next.
This suggests a funny mismatch: the AI learned the dance steps (who moves when) perfectly, but it hasn't quite learned the music (how to say things in a way that feels human). The study suggests that understanding the social structure of a group is a huge step forward, but it doesn't automatically make the AI a better conversationalist. It's like teaching a robot the rules of soccer so well that it knows exactly when to pass the ball, but when it actually kicks the ball, it still trips over its own feet. The researchers conclude that while we are getting better at modeling the flow of group talk, we still have a long way to go before AI can truly join a human team and chat like one of the gang.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.