← Latest papers
💻 computer science

Human-robot conversation with multiple participants in noisy public spaces

This paper presents a multi-channel microphone array audio system designed for noisy public spaces that enhances speech for both attentive listening by the android ERICA and remote avatar interactions via Teleco robots, while also providing spatial audio to improve immersion in multi-party conversations.

Original authors: Divesh Lala, Yogeeswaran Muthukumaran, Vincent Fernandes, Kazushi Kato, Shota Fujiki, Zihao Chi, Masaya Iwasaki, Taiken Shintani, Megumi Kawata, Kazuki Sakai, Koji Inoue, Yuicihiro Yoshikawa, Tatsuya
Published 2026-09-02
📖 6 min read🧠 Deep dive

Original authors: Divesh Lala, Yogeeswaran Muthukumaran, Vincent Fernandes, Kazushi Kato, Shota Fujiki, Zihao Chi, Masaya Iwasaki, Taiken Shintani, Megumi Kawata, Kazuki Sakai, Koji Inoue, Yuicihiro Yoshikawa, Tatsuya Kawahara

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to have a serious conversation in the middle of a bustling train station, where the roar of engines, the chatter of crowds, and the echo of the architecture compete for your attention. Humans manage this feat effortlessly, a skill scientists call the "cocktail party effect," where our brains filter out the background din to focus on a single voice. For a robot, however, this is a formidable challenge. While artificial intelligence has become remarkably good at holding text-based conversations, giving a robot the ability to listen and speak naturally in a noisy, crowded public space remains a significant hurdle. This is especially true when multiple people are talking at once, or when the robot itself is moving through the environment. Without a way to isolate human speech from the chaos, a robot's ability to understand and respond collapses, leaving it deaf to the very people it is meant to serve.

Researchers from several universities in Japan and Singapore set out to solve this problem by building a system that allows robots to hear clearly in the real world. They tested their work at the 2025 World Expo in Osaka, a setting chosen specifically because it is loud, unpredictable, and filled with thousands of visitors. The team developed a specialized audio system capable of capturing the voices of multiple people simultaneously, even when they are speaking over one another, and cleaning up the sound so that both the robot and a remote human operator can understand them. The result was a demonstration where robots successfully held conversations with groups of people in a noisy pavilion, proving that machines can be taught to listen with the same focus a human uses in a crowded room.

The core of this achievement lies in how the robots are equipped to hear. Instead of relying on a single microphone, which would pick up all the noise in the room equally, the researchers used a circular array of four microphones. This device acts like a set of ears that can be electronically steered to focus on specific directions. When a person speaks, the system identifies their location and enhances their voice while suppressing the background noise coming from other directions. This process creates a clean audio signal that can be used for two distinct purposes: it allows the robot's internal computer to recognize the words being spoken, and it sends a clear, high-quality version of the conversation to a human operator who might be controlling the robot from a distant location.

The team demonstrated this technology in two different scenarios, each with its own unique challenges. In the first setup, an android robot named ERICA sat with two human participants. The goal was to show that the robot could act as an attentive listener, following the conversation and responding with nods or short verbal cues. Because the humans were seated in a fixed position, the microphone array could be placed on a tripod between them, creating a stable listening environment. The system successfully separated the voices of the two people, allowing ERICA to understand who was speaking and to generate appropriate responses, such as asking follow-up questions based on what was just said. This scenario proved that the technology could handle the complexities of a group conversation where people might interrupt or speak at the same time.

The second scenario was far more difficult because it involved movement. Here, the researchers used small, mobile robots called Telecos. One of these robots was controlled by a human operator sitting in a separate room, acting as a virtual avatar, while the other robot moved around on its own to support the conversation. In this case, the microphone array was mounted on the head of the robot controlled by the human. As the robot turned its head or moved across the floor, the direction of its "ears" changed constantly. The researchers had to program the system to continuously track the position of the human participant and the other robot, instantly switching the microphone's focus to whichever direction the conversation was coming from. This dynamic adjustment allowed the remote operator to hear the human participant clearly, even as the robot moved, and to feel as though they were physically present in the conversation.

To make this work, the system had to do more than just amplify sound; it had to create a sense of space. When the remote operator listened through headphones, the audio was adjusted so that if the human participant was standing to the left of the robot, the voice would come from the left side of the operator's headphones. This spatial audio effect helped the operator feel immersed in the interaction, preserving the natural relationship between the speakers. Additionally, the system included an echo cancellation feature to prevent the operator's own voice, which was being played through the robot's speaker, from being picked up by the microphone and confusing the system.

The results of the demonstration were encouraging. Over six days at the Expo, the systems engaged with hundreds of visitors in a pavilion filled with the noise of other exhibits and the occasional sound of rain. The robots were able to maintain coherent conversations, and the remote operators reported that they could clearly hear the people they were talking to, despite the high noise levels. In controlled tests prior to the event, the system's ability to recognize speech was comparable to using a standard handheld microphone, even when the robot was moving or when there was significant background noise. The researchers noted that while the system was not perfect—sometimes it missed quiet speakers or struggled when a person turned their back to the microphone—it represented a significant step forward in making robots capable of functioning in real-world public spaces.

This work highlights a crucial shift in how we think about human-robot interaction. It moves beyond the controlled, quiet environments of a laboratory and into the messy, loud reality of daily life. By solving the problem of how to hear clearly in a crowd, the researchers have paved the way for robots that can serve as companions, guides, or assistants in places like airports, museums, and shopping centers. The ability to filter out the noise and focus on the human voice is not just a technical feat; it is the foundation for building machines that can truly understand and connect with us, even in the most chaotic of settings.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →