← Latest papers
💬 NLP

On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous Levels

This paper challenges the traditional view of addressee detection as a discrete classification task by proposing a continuous framework that better captures the graded nature of address and demonstrates its superior predictive power for turn-taking and listener behaviors like gaze and backchannels.

Original authors: Taiga Mori, Koji Inoue, Divesh Lala, Tatsuya Kawahara

Published 2026-07-20
📖 5 min read🧠 Deep dive

Original authors: Taiga Mori, Koji Inoue, Divesh Lala, Tatsuya Kawahara

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking into a bustling coffee shop where three friends are chatting with a very polite, but slightly confused, robot. In the world of computer science, specifically in a field called "Natural Language Processing," researchers try to teach machines to understand human conversation. One of the trickiest puzzles they face is figuring out who a person is talking to. Is the speaker talking to the robot? To their friend on the left? To the friend on the right? Or are they addressing the whole group at once?

For a long time, scientists treated this like a simple game of "Pick One." They assumed that every sentence had to be aimed at exactly one target, like throwing a ball that can only land in one specific basket. If you were talking to the whole group, the computer just had to pick one label for that. But this feels a bit like trying to describe a sunset using only the words "orange" or "blue." Real conversations are messy, fluid, and often feel like they are being sent to multiple people at different strengths. Understanding who is being talked to is crucial because it helps computers know when to jump in and speak, or when to just listen and nod. If a robot thinks you are talking to your friend when you are actually asking it a question, it might stay silent when it should answer, or worse, interrupt your friend.

This brings us to a new study by Taiga Mori and his team from Kyoto University. They decided to stop treating "who is being talked to" as a simple on/off switch and started looking at it as a dimmer switch. Instead of asking, "Is this sentence for Person A or not?" they asked, "How much is this sentence for Person A, and how much for Person B?"

To find the answer, the researchers listened to 36 recorded conversations between groups of three Japanese friends. They didn't just ask one person to guess who was being talked to; they had three different people listen to every single sentence and write down their guesses. They found that the guessers often disagreed. Sometimes one person thought a sentence was for the whole group, while another thought it was just for one friend. The old way of doing things would call this "noise" or a mistake. But the researchers suspected this disagreement wasn't an error; it was actually a clue that the feeling of being "addressed" is naturally fuzzy and continuous.

So, they built a special mathematical model to turn those fuzzy human guesses into a "continuous address level." Imagine a slider that goes from 0 to 1. If a speaker is talking directly to you, the slider is at 1.0. If they are ignoring you, it's at 0.0. But what if they are talking mostly to you, but kind of including your friend too? The slider might sit at 0.7. This new "slider" system allowed them to measure the strength of the connection between the speaker and each listener.

Then, they tested if this new "slider" idea worked better than the old "yes/no" labels. They looked at three things: who spoke next, where the listeners were looking, and how often they made little sounds like "uh-huh" or "yeah" (called backchannels).

The results were pretty clear. When they tried to predict who would speak next, the model using the "slider" (continuous levels) was much better at guessing than the model using the simple "yes/no" labels. It turns out that even if a listener isn't the main target of a sentence, if they are being addressed somewhat, they are more likely to jump in and speak. The old "yes/no" system missed these subtle chances to speak.

The same thing happened with eye contact. The study found that the more a listener felt addressed (the higher the slider was), the more they looked at the speaker. But it wasn't a simple switch where you either look or you don't. The amount of eye contact changed gradually, like a volume knob turning up, matching the strength of the address.

Finally, they looked at those little "uh-huh" sounds. Again, the "slider" model did a better job predicting them. The more a listener felt included in the conversation, the more they tended to make those supportive noises. However, the improvement here wasn't as huge as it was for speaking or eye contact, suggesting that while being addressed matters, other things also influence whether someone makes a sound.

The authors suggest that this means we need to rethink how we build conversation robots. Instead of forcing a computer to decide, "I am talking to Person A," it should be able to say, "I am talking mostly to Person A, but also a little bit to Person B." This continuous view captures the messy, beautiful reality of human chat much better than the rigid boxes we used to use. It suggests that in a group conversation, a speaker isn't just picking one person to talk to; they are distributing their attention like a spotlight that can shine brightly on one person, dimly on another, or wash over the whole room all at once.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →