Sound-based Multi-Person 3D Pose Estimation
This paper introduces SoundMHPE, the first framework to estimate multi-person 3D poses solely from acoustic signals by utilizing a novel encoder-decoder architecture to overcome challenges like signal superposition and reflections, validated on a newly constructed 6-hour synchronized dataset.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
For decades, scientists have sought ways to see the invisible, developing technologies that can track human movement without relying on the human eye. Cameras, which depend on light, fail in the dark or when a person is hidden behind a wall. Radio waves, used in some sensing systems, struggle when they encounter water or metal. In these complex, real-world scenarios, a different kind of signal has emerged as a promising alternative: sound. By actively sending out a sound pulse and listening carefully to how it bounces back, researchers can map the shape and position of objects in a room. This technique, known as active acoustic sensing, has already shown it can determine the posture of a single person, even when they are obscured. However, a significant gap remained in this field. While a single voice or movement creates a distinct echo, a room filled with multiple people creates a chaotic mix of overlapping sounds. Untangling these signals to understand where each individual is standing and how they are moving has been considered nearly impossible, as the echoes from one person blend indistinguishably with those of another.
A team of researchers from Keio University, the Tokyo University of Science, and NTT has now taken the first step toward solving this problem. They have developed a system capable of estimating the three-dimensional poses of multiple people using only sound. The researchers built a custom dataset to train their system, recording six hours of synchronized audio and motion data from up to three people moving simultaneously in a room. The setup involved a pair of speakers emitting a specific sound signal and a specialized microphone capable of capturing sound from all directions. As the participants walked, twisted, and raised their arms, the system recorded the returning echoes. The challenge was immense; the acoustic signals were a superposition of many different movements, further complicated by the way sound reflects off bodies and walls, creating delays that obscure the direct link between a specific pose and a specific sound change.
To navigate this complexity, the team created a new framework they call SoundMHPE. Instead of trying to process the sound as a single, messy block, their system breaks the audio down into multiple layers of detail. It analyzes the sound using different time windows, looking at both rapid changes that happen in a split second and slower, finer details in the frequency of the sound. This allows the system to isolate subtle acoustic signatures that might otherwise be lost in the noise. Once the sound is analyzed, the system uses a sophisticated method to separate the people. It assigns a specific set of digital "queries" to each individual, allowing the software to track the temporal evolution of one person's movements while simultaneously understanding how they interact with the others in the room. This approach disentangles the overlapping information, enabling the system to reconstruct the pose of each person frame by frame.
The results of this work demonstrate that the system works. When tested against existing methods that were originally designed for single-person tracking or adapted from radio-wave sensing, the new sound-based system proved more accurate. It successfully tracked complex movements, such as a person twisting their body or raising both arms, even when performed alongside other moving individuals. In tests involving two and three people, the system maintained a high level of precision, outperforming the baseline models in measuring the distance between predicted and actual joint positions. The researchers also found that their method could adapt to unseen environments; when they placed partitions in the room to change how sound reflected, the system still managed to estimate coarse poses, suggesting it can handle different acoustic conditions. Furthermore, the underlying logic of their system was shown to be effective even when applied to radio frequency data, indicating that the approach is robust across different types of signals.
This research marks a significant expansion of what is possible with acoustic sensing. By moving from single-person scenarios to multi-person environments, the team has opened the door for applications in crowded spaces, disaster relief, and sports analysis where cameras cannot be used and radio signals might be blocked. While the technology is still in its early stages and requires further development to be deployed in the real world, the construction of this new dataset and the validation of the multi-person estimation method provide a crucial foundation. The work confirms that sound, when analyzed with the right tools, can reveal the hidden geometry of a group of people, turning a chaotic mix of echoes into a clear picture of human movement.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.