← Latest papers
💻 computer science

PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic Interaction

This paper introduces PolySLGen, an online framework that generates contextually appropriate and temporally coherent multimodal speaking and listening reactions for target participants in polyadic group interactions by effectively aggregating motion and social cues from all group members.

Original authors: Zhi-Yi Lin, Thomas Markhorst, Jouh Yeong Chew, Xucong Zhang

Published 2026-04-10
📖 5 min read🧠 Deep dive

Original authors: Zhi-Yi Lin, Thomas Markhorst, Jouh Yeong Chew, Xucong Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking into a lively dinner party with four friends. You aren't just talking to one person; you are part of a group. To fit in, you have to do three things at once:

  1. Listen to what everyone is saying.
  2. Watch who is looking at whom and how they are moving their hands.
  3. React instantly—either by jumping in with a joke (speaking) or by nodding and smiling (listening).

Doing this naturally is incredibly hard for robots or AI. Most current AI is like a person at a party who can only talk to one person at a time, or worse, they only know how to speak and forget how to listen. They might stand there awkwardly staring at the wall when they aren't talking.

Enter PolySLGen. Think of it as the "Ultimate Party Guest" AI.

The Problem: The "Two-Person" Trap

Most AI researchers have been training robots to have conversations like a phone call between two people (dyadic). But real life is a group chat (polyadic).

  • The Old Way: If you asked an old AI to join a group, it would get confused. It might ignore the person sitting next to you because it's only looking at the person it's "talking" to. It also struggles to know when to speak and when to shut up and listen.
  • The Missing Piece: It doesn't just need to hear words; it needs to see body language. If your friend is looking at you while speaking, they want you to reply. If they are looking at someone else, you should stay quiet.

The Solution: PolySLGen (The "Social Brain")

The researchers built a system called PolySLGen (Polyadic Speaking-Listening Generation). Here is how it works, using a simple analogy:

1. The "Group Hug" Sensor (Pose Fusion)

Imagine you are in a room with four people. Instead of looking at them one by one, PolySLGen has a special "group hug" sensor. It takes the movements of everyone in the room and blends them into a single, compact "vibe" signal.

  • Why? If the AI tried to process every single person's hand movement separately, it would get overwhelmed (like trying to listen to four radio stations at once). By fusing them, it understands the group dynamic instantly.

2. The "Eye Contact" Detector (Social Cue Encoder)

Have you ever been in a group and felt like someone was ignoring you? That's because they aren't looking at you.
PolySLGen has a special module that tracks head orientation. It asks: "Who is looking at our target robot?"

  • If everyone is looking at the robot, the AI knows: "Okay, it's my turn to speak!"
  • If the robot is looking at someone else, the AI knows: "I should be listening and nodding."
    This is crucial for knowing when to switch from "Listening Mode" to "Speaking Mode."

3. The "Dual-Mode" Actor (Speaking & Listening)

Old AI models usually have two separate brains: one for talking and one for moving. PolySLGen has one brain that handles both.

  • Speaking Mode: It generates words, a voice, and gestures that match the words.
  • Listening Mode: It doesn't just freeze! It generates natural reactions like nodding, tilting the head, or shifting weight, exactly like a human does when they are paying attention.
  • The Switch: It predicts a "Speaking Score." If the score is high, it talks. If it's low, it listens. This makes the transition smooth, like a real conversation, rather than a robotic "ON/OFF" switch.

How They Tested It

They tested this AI in a Dungeons & Dragons (D&D) game setting.

  • The Scenario: A group of players and a "Dungeon Master" (the target) are playing a game. They are shouting, laughing, and moving around a table.
  • The Test: They asked the AI to be the Dungeon Master.
  • The Result:
    • Old AI (SOLAMI): Would often freeze when it wasn't talking, or make weird movements that didn't match the conversation. It looked like a robot trying to act human.
    • PolySLGen: It looked natural. It knew when to interrupt, when to wait, and how to gesture while listening. In a "blind test" where humans watched videos, they preferred PolySLGen every time because it felt "real."

Why This Matters

This isn't just about making a cooler robot. It's about making AI that can actually hang out with us.

  • Education: A robot teacher that can manage a whole classroom, not just one student.
  • Therapy: A robot therapist that can sit with a group of people and understand the group's emotional flow.
  • Customer Service: A robot receptionist that can greet a family of four, not just the person at the front.

The Bottom Line

PolySLGen is like teaching an AI the art of "reading the room." It doesn't just listen to words; it watches the whole group, tracks eye contact, and knows exactly when to speak up and when to sit back and nod. It's a giant leap from "robot that talks" to "robot that hangs out."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →