Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction
This paper presents a real-time human-centric world model that synthesizes coherent upper-body interactions by combining a unified multi-scale implicit representation for continuous human motion with language-encoded discrete states for precise object contact control, achieving 25 FPS inference on H100 GPUs.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a movie where the main character can reach out, grab a coffee cup, take a sip, and set it down, all while talking to you. For a long time, computer scientists have been great at making the character's body move realistically, but they often struggled with the "grabbing" part. Sometimes the hand would float through the cup like a ghost, or the cup would magically change shape when touched. This is the challenge of "human-object interaction": teaching a computer to understand not just how a person moves, but how they physically touch and change the world around them.
To do this, researchers usually rely on two main tools. First, there's video generation, which is like a digital artist that can paint moving pictures based on a description. Second, there's world modeling, which is a fancy way of saying a computer tries to predict how a scene will change over time. The big question has been: Can we combine these to make a character that doesn't just dance in place, but actually interacts with objects in real-time, like a video game character that feels real enough to touch?
This is exactly what a team of researchers from Tongyi Lab has been working on. They built a new system that acts like a "real-time human-centric world model." Think of it as a digital puppet master that doesn't just control the puppet's limbs, but also decides if the puppet is holding a hammer or just waving its hand in the air. Their goal was to create a system that can take a live video of a person, a picture of an object (like a cup), and a simple command (like "grasp" or "no contact"), and instantly generate a video where the person interacts with that object perfectly.
The team found that by splitting the control into two distinct parts, they could solve the problem much better than previous methods. The first part is the continuous human state, which handles the smooth, flowing movements of the body, hands, and face. Imagine this as the puppeteer's hand guiding the puppet's muscles. The second part is the discrete interaction state, which is like a simple switch that tells the system, "Is the hand touching the object right now?" or "Is it holding it?" By using a specific, short command for this switch (like the words "grasp" or "no contact") instead of a long, complicated sentence, the computer can switch between these states instantly and accurately.
The result is a system that can generate these interactions at a speed of 25 frames per second with a delay of only about 1000 milliseconds. This means it's fast enough to feel like a live conversation or a real-time video game. The researchers showed that their method creates more realistic hand movements and better contact with objects compared to older systems, which often produced weird glitches like floating objects or hands that didn't quite close around what they were supposed to hold. They also created a special dataset of over 30,000 examples to teach their system how these interactions should look. While they haven't claimed this solves every problem in the universe, their experiments suggest that this "continuous-discrete" approach is a significant step forward in making digital humans that can truly interact with their world in real-time.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.