← Latest papers
🤖 machine learning

VSMP-IMU: Video-Grounded Semantic Motion Programs for Sensor-Aware Synthetic IMU Generation

The paper introduces VSMP-IMU, a video-grounded framework that generates controllable synthetic IMU data using structured Semantic Motion Programs to overcome data scarcity and significantly improve wearable human activity recognition performance across low-resource, imbalanced, and subject-generalization scenarios.

Original authors: Lala Shakti Swarup Ray, Vitor Fortes Rey, Mengxi Liu, Paul Lukowicz, Bo Zhou

Published 2026-08-07
📖 4 min read☕ Coffee break read

Original authors: Lala Shakti Swarup Ray, Vitor Fortes Rey, Mengxi Liu, Paul Lukowicz, Bo Zhou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to recognize when you are dancing, running, or doing yoga. To do this, the robot needs to "feel" your movements, usually through a smartwatch or a sensor strapped to your body that records how you shake and spin. This field is called Human Activity Recognition. The problem is that teaching a robot requires thousands of examples of real people moving, and getting those examples is a nightmare. You have to strap sensors on real humans, ask them to repeat the same move over and over, and hope they don't get tired or change how they do it. It's expensive, slow, and often leaves the robot confused when it meets a new person who moves slightly differently.

To solve this, scientists have tried to "fake" the sensor data. Some try to watch a video of a person and guess what the sensor would feel (like looking at a dance video and guessing the rhythm). Others try to write a text description like "jump up and down" and have a computer imagine the movement. But these methods have flaws: the video method gets confused if the camera angle is weird, and the text method is too vague, often creating movements that sound right but feel wrong to a sensor. The big question is: can we create fake sensor data that is both realistic enough to teach the robot and flexible enough to cover all the different ways humans actually move?

Enter VSMP-IMU, a new system that acts like a super-smart director for a movie about movement. Instead of just guessing from a video or a text prompt, this system uses a "Semantic Motion Program" (SMP). Think of an SMP as a detailed recipe card for a movement. It doesn't just say "jumping jacks"; it breaks the move down into specific steps: "stand, jump open, close, repeat," and notes exactly which body parts are involved, how fast it should go, and how wide the arms should swing.

Here is how the magic happens: The system watches a video of a real person doing an activity and extracts this "recipe card." Then, it gets creative. It can tweak the recipe—maybe ask for the jumping jacks to be done super fast or with huge arm swings—while keeping the core identity of the move (it's still a jumping jack). Once the recipe is tweaked, a computer generates a 3D animation of a person doing exactly that. Finally, the system simulates what a sensor strapped to that animated person's body would feel, turning the 3D motion into fake sensor data.

The researchers tested this by feeding this new, fake data into a robot's brain to see if it got better at recognizing real human movements. The results were impressive. When the robot was trained only on real data, it got a score of about 68.6 out of 100. When they added the VSMP-IMU fake data, the score jumped to 78.3. That's a huge improvement, especially when there isn't much real data to begin with. In situations where the robot had very little real data to learn from (like only 1% of the usual amount), the system helped the robot improve its score by nearly 19 points compared to learning from real data alone.

The system also proved it could follow instructions. When the researchers asked the system to generate movements with a specific speed, the resulting motion matched the request 92% of the time. When they asked for a specific number of repetitions, the system got it right within one count 90.6% of the time. However, the paper notes that while the system is great at big, whole-body movements like jumping or squatting, it sometimes struggles with tiny, precise hand movements (like catching a ball) because the underlying 3D animation doesn't show every single finger twitch.

In short, VSMP-IMU suggests that by using a structured "recipe" to guide the creation of fake sensor data, we can teach robots to understand human movement much better than before. It bridges the gap between watching a video and feeling a vibration, creating a library of synthetic movements that are diverse, controllable, and surprisingly realistic. While it's not perfect for every tiny detail of human motion, it offers a powerful new way to train the next generation of wearable technology without needing to strap sensors on thousands of real people.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →