COMPASS: Complete Multimodal Fusion via Proxy Tokens and Shared Spaces for Ubiquitous Sensing
The paper proposes COMPASS, a robust multimodal fusion framework that addresses missing modality challenges by synthesizing target-specific proxy tokens in a shared latent space to ensure the fusion head always receives a fixed, complete input structure, thereby outperforming existing methods across diverse sensing scenarios.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a complex puzzle, like figuring out what a person is doing in a room (walking, dancing, falling). To do this, you have a team of five different detectives, each with a special skill:
- The Camera (sees the visual scene).
- The Radar (sees movement through walls).
- The WiFi Sensor (senses how radio waves bounce off bodies).
- The RFID Tag (tracks specific devices).
- The Depth Sensor (measures distance).
In a perfect world, all five detectives are in the room, and they all whisper their findings to a Chief Detective (the "Fusion Head"), who combines their stories to give the final answer.
The Problem: The Missing Detective
In the real world, things go wrong. Maybe the camera is broken, or the WiFi signal is weak, or someone is hiding behind a wall. Suddenly, one or more detectives are missing.
Old methods tried to handle this in two ways:
- The "Skip" Method: If the Camera is missing, the Chief Detective just ignores that spot and tries to solve the puzzle with only four voices. But the Chief was trained to listen to five voices at once. When one is missing, the conversation gets awkward, the logic breaks, and the answer is often wrong.
- The "Imitation" Method: They try to guess what the missing detective might have said based on the others. But often, this guess is just a vague approximation that doesn't quite fit the Chief's specific way of listening.
The Solution: COMPASS (The "Stand-in" System)
The paper introduces COMPASS, a clever new system that changes the rules of the game. Instead of letting the Chief Detective deal with a missing voice, COMPASS ensures the Chief always hears exactly five voices, no matter what.
Here is how it works, using a simple analogy:
1. The "Proxy" Stand-ins
When a detective (say, the Camera) is missing, COMPASS doesn't leave the seat empty. Instead, it instantly creates a Proxy (a stand-in actor) to sit in that seat.
- This isn't a random guess. The other detectives (Radar, WiFi, etc.) talk to a special "Translator" who knows exactly how to turn their observations into a story that sounds like the missing Camera would have told it.
- If the Radar sees a person moving fast, the Translator creates a "Camera Proxy" that says, "I see a person moving fast," even though the Camera is broken.
2. The "Shared Language" Room
For this to work, all the detectives must speak a Shared Language.
- Normally, a Radar speaks in "waves" and a Camera speaks in "pixels." They are hard to translate.
- COMPASS forces everyone to translate their raw data into a common "secret code" (a shared latent space) before they talk. This ensures that when the Radar says "moving fast," the Proxy knows exactly how to write that down in the Camera's language.
3. The "Fixed Table"
The most important rule of COMPASS is: The table always has five chairs.
- Whether all five detectives are there, or only one is there, the Chief Detective sits at a table with exactly five slots.
- If a detective is present, they sit in their chair.
- If they are missing, their Proxy sits in their chair.
- The Chief never has to change how they listen. They just listen to the five voices in front of them. This makes the system incredibly robust and fast.
Why is this better?
- No Confusion: The Chief Detective doesn't have to relearn how to listen every time a detective goes on break. The input is always the same.
- Smart Stand-ins: The Proxies aren't just random noise; they are trained to be useful. They are taught to not only sound like the missing detective but also to help solve the specific puzzle (e.g., "Is this person falling?").
- Speed: Because the Chief doesn't have to do complex math to figure out "who is missing," the system runs much faster than previous methods.
The Results
The researchers tested this on three different "crime scenes" (datasets) involving human activity recognition.
- In scenarios where sensors failed or were blocked, COMPASS was significantly better than the old methods.
- For example, if only the WiFi sensor was working, old methods struggled. But COMPASS used the WiFi sensor to create a "Radar Proxy" and a "Camera Proxy," allowing the system to guess the missing details with high accuracy.
In Summary
COMPASS is like a super-organized meeting where, if someone is absent, a highly trained substitute immediately takes their seat and speaks in their voice. This way, the meeting (the AI model) can always run smoothly, regardless of who shows up, ensuring the final decision is always accurate.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.