Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots
This paper introduces AffectLoop, a multimodal empathetic social robot system that closes the affective loop by dynamically tracking both speaker and listener emotional states to generate congruent verbal and embodied responses, which a pilot study shows significantly improves user satisfaction and distress recovery compared to text-only baselines.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Human connection relies on more than just the exchange of words. When we speak with someone, we are constantly reading their face, listening to the tone of their voice, and feeling the rhythm of the conversation. We also bring our own mood into the interaction, which shifts and changes as we listen and respond. This dynamic flow of emotion is what makes a conversation feel alive. For decades, researchers have tried to teach machines to understand this complexity, building systems that can recognize when a person is sad or happy. However, most of these digital listeners operate like a one-way street: they detect a user's emotion and offer a pre-programmed reply, ignoring the fact that the listener's own reaction also changes the emotional landscape of the conversation. They miss the loop where two people influence each other's feelings in real time.
A team of researchers has now built a robot that attempts to close this gap. In a study presented at the 2026 Asia Pacific Signal and Information Processing Association Annual Summit, they introduced a system called AFFECTLOOP, designed to help a social robot understand not just what a person says, but how the emotional exchange evolves between them. The researchers equipped a small, programmable robot named Misty II with the ability to track the speaker's voice and facial expressions while simultaneously monitoring its own internal emotional state. By feeding this dual stream of information into a large language model, the robot generates responses that are not only empathetic but also physically expressive, creating a continuous cycle of emotional connection.
The core idea behind AFFECTLOOP is that true empathy requires a two-way street. In a typical conversation, a person's emotions are not static; they shift from moment to moment based on what is being said and how the other person is reacting. The researchers found that existing robotic systems often treat emotion as a single label attached to a sentence, missing the subtle dynamics of how feelings change over time. To fix this, their system breaks down the interaction into two parallel tracks. First, it analyzes the human speaker, using software to convert their voice into a text transcript and their facial movements into a series of emotional signals. It tracks the speaker's mood as it rises and falls, noting changes in their energy and emotional intensity. Second, the system keeps a running record of the robot's own state. It calculates the emotional tone of the words it plans to say and the feelings conveyed by its physical movements, such as head tilts or gestures.
These two streams of data are combined to guide the robot's next move. The robot does not simply react to the human; it considers how its own previous actions and current mood fit into the ongoing conversation. This information is passed to an artificial intelligence engine, which acts as the robot's brain. The engine is instructed to act as an attentive listener, generating a short, spoken response that matches the emotional context. Crucially, the engine also decides on a physical action for the robot to perform, ensuring that its body language aligns with its words. For example, if the human is sharing a sad story and the robot detects a shift toward distress, it might generate a gentle verbal comfort while simultaneously lowering its head or softening its gaze. Once the robot speaks and moves, the system immediately re-evaluates the situation, updating its own emotional state to reflect what just happened, and then waits for the human's next turn. This creates a closed loop where the robot's reaction becomes part of the emotional dynamic, just as it would in a human conversation.
To test whether this approach actually worked, the researchers conducted a small study with five participants. Each person spent five minutes talking to the robot in two different scenarios. In the first scenario, the robot used a standard system that only looked at the human's words to generate a response, ignoring the robot's own emotional state and the dynamic flow of the conversation. In the second scenario, the robot used the new AFFECTLOOP system. Afterward, the participants rated their experience on several factors, including how natural the conversation felt, how well the robot seemed to listen, and how supportive the robot's responses were.
The results suggested that the new system made a noticeable difference. Participants rated the AFFECTLOOP robot higher overall, with the most significant improvement seen in how supportive and empathetic the robot's responses felt. People felt more encouraged and praised by the robot when it used the closed-loop system. They also reported feeling better after the conversation, with their stress levels appearing to decrease more effectively than with the standard system. The researchers also analyzed the logs of the interactions to see what was happening beneath the surface. They found that the robot using AFFECTLOOP was better at aligning its emotional shifts with the human's. When the human's mood moved in a certain direction, the robot's reaction tended to follow that same path more closely, rather than just reacting randomly. Furthermore, the data showed that when the human started the conversation feeling negative, they were more likely to end up in a less negative state after talking to the AFFECTLOOP robot.
While the study involved a small number of participants and serves as a preliminary test, the findings point toward a promising direction for social robotics. The researchers suggest that by explicitly modeling both the speaker's changing emotions and the listener's own affective state, robots can move beyond simple text-based responses to create a more genuine, embodied form of empathy. The system did not make the robot seem more human-like in every way; in some measures, the standard system still scored slightly higher on perceived naturalness. However, the clear gains in emotional support and user satisfaction indicate that closing the affective loop helps the robot feel more like a supportive companion. This work demonstrates that for a robot to truly connect with a human, it must not only listen to the other person but also be aware of its own place in the emotional dance of the conversation.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.