← Latest papers
💬 NLP

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

This paper investigates how gestures and speech jointly convey referential information in video-mediated dialogue under varying visibility conditions, demonstrating that gesture alone is predictive, multimodal fusion is most beneficial when speech is ambiguous, and partner visibility influences both gesture production and interactional entrainment.

Original authors: Esam Ghaleb, Hugh Mee Wong, Kristina Kobrock

Published 2026-08-11
📖 3 min read☕ Coffee break read

Original authors: Esam Ghaleb, Hugh Mee Wong, Kristina Kobrock

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to guess a secret object your friend is thinking of, but you can only hear their voice. Sometimes they say, "It's the one with the pointy bit," which is pretty vague. But if you can also see them, they might point at the air or wiggle their fingers to show the shape. This is the magic of human conversation: we don't just use words; we use our whole bodies to fill in the gaps. Scientists who study how computers understand language (a field called Natural Language Processing) have been great at reading transcripts of what people say, but they often struggle to understand the messy, unspoken parts of a conversation, like hand gestures. The big question is: Can a computer learn to "read" those hand movements to figure out what someone is talking about, especially when the words alone aren't clear enough? And does it matter if the computer can "see" the person making the gestures, or just hear them?

This paper dives into that exact mystery by setting up a digital game. The researchers created a scenario where two people had to identify one of 16 strange, alien-like shapes (called "Fribbles") that have no real names. They had to describe and pick the right one out of a lineup. The team tested this in three different "visibility" settings: where the partners could only hear each other, where they could see each other from one side, and where they could see each other fully. They built computer models to act as the guesser, feeding them either just the spoken words, just the skeleton-like movement of the hands, or a mix of both.

Here is what they found: The hand gestures alone were actually surprisingly good at guessing the right object. The computer models could pick the correct Fribble just by watching the hand movements about 20% of the time, which is way better than random guessing (which would only get it right 5.9% of the time). However, the words were still the star of the show, getting it right about 46.5% of the time. The real magic happened when the computer combined both. The hand gestures didn't just add a little bit of help; they acted like a safety net. When the computer was confused by the words and didn't know what to guess, the hand movements provided the extra clue needed to solve the puzzle.

The researchers also discovered that the "visibility" of the partners changed how the humans acted. When the partners could see each other, they used more gestures, and those gestures were packed with more useful information. But when they couldn't see each other, the gestures became a bit more like self-soothing movements and less helpful for the computer to guess the object. Interestingly, as the partners played the game over and over again, they got better at describing the objects with words, but their hand gestures didn't get any better or more specific over time.

Finally, the team tried a fancy trick to help the computer understand gestures better. They taught the model to mentally "match" the shape of a hand movement to the picture of the object it was describing. This didn't make the hand gestures better on their own, but it did make the combination of words and gestures even sharper, pushing the success rate up a few more percentage points. In short, the paper suggests that while computers are getting good at listening, they still need to learn how to watch us move to truly understand what we mean, especially when our words get a little fuzzy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →