← Latest papers
🤖 AI

HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction

HiPHI is a large-scale, high-precision benchmark dataset comprising over 600 hours of whole-body human motion and object-interaction data, designed to overcome the limitations of existing embodied datasets by combining sub-millimeter accuracy with theoretically guided diversity to enable scalable training and evaluation of humanoid policies.

Original authors: Jiahao Ji, Ji Ma, Runhan Zhang, Runyi Yu, Wenjia Wang, Weiheng Chi, Qianqian Peng, Weichao Yan, Yongfei Gu, Ye Tian, Ting Wu, Longwei Li, Chun Yuan, Ruoli Dai, Lei Han

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Jiahao Ji, Ji Ma, Runhan Zhang, Runyi Yu, Wenjia Wang, Weiheng Chi, Qianqian Peng, Weichao Yan, Yongfei Gu, Ye Tian, Ting Wu, Longwei Li, Chun Yuan, Ruoli Dai, Lei Han

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

To teach a robot to move like a human, engineers must first solve a paradox. They need data that is broad enough to cover the vast variety of things a person might do, yet precise enough to capture the exact physics of how a foot lands or a hand grips a handle. For years, the field has been stuck between two imperfect options. On one side, there are millions of hours of internet videos. These show a huge range of human behavior, but they are just pictures; they lack the hidden physical details, like the exact weight of an object being lifted or the precise force of a step. On the other side, there are laboratory recordings made with special cameras that track every joint of a body. These are incredibly accurate, but they usually cover only a narrow set of actions, often missing the messy, complex interactions with real objects that happen in the real world. Without a bridge between these two worlds, robots struggle to learn how to move safely and effectively in a physical environment.

A team of researchers has built that bridge with a new dataset called HiPHI. They created a massive library of human movement, capturing more than 600 hours of high-fidelity motion from 132 different people. Unlike previous attempts that relied on writers manually inventing lists of actions to film, this team used a structured system based on how language organizes events. They started with a linguistic framework that breaks down human actions into specific, meaningful units, then systematically expanded each unit by changing factors like speed, direction, and the objects being used. The result is a dataset that includes not just the movement of the human body, but also the synchronized movement of real objects, from furniture to sports equipment. By recording the body and the object together with sub-millimeter accuracy, the researchers ensured that the data reflects the true physical constraints of the world, such as friction, weight, and balance.

The collection process was designed to be exhaustive rather than accidental. Instead of filming actors performing random scripts, the team treated the creation of the dataset like building a map. They began with a set of core action concepts, such as "walking" or "carrying," and then generated hundreds of variations for each one. A single concept like walking was filmed at different speeds, in different directions, with different stride lengths, and while carrying loads of varying weights. This approach allowed them to cover the entire space of possible human motions in a logical way, ensuring that no major type of movement was left out. The final dataset contains over 200 million frames of motion, including nearly 250 hours of sequences where people interact with 40 distinct real-world objects. Every clip is tagged with a specific label describing the action and the context, making it easy for researchers to find exactly the type of movement they need to teach a robot.

The quality of this data was tested rigorously to ensure it could actually be used to train robots. The researchers checked for common errors that plague other datasets, such as feet sinking into the floor or bodies floating in mid-air without support. Their measurements showed that HiPHI is significantly cleaner and more physically consistent than existing large-scale motion libraries. When they used this data to train a humanoid robot in a computer simulation, the robot learned to walk, sit, and carry objects much faster and more reliably than when trained on other datasets. The robot was able to maintain its balance and execute complex tasks, such as pushing a heavy box or flipping a suitcase, with a level of stability that was previously difficult to achieve. These simulations suggested that the data captures the subtle physical cues necessary for a machine to understand how to move its own body in relation to the world around it.

The success of the project was further confirmed when the team took the trained robot out of the simulation and onto the floor of a real laboratory. Despite the challenges of real-world sensors and mechanical limits, the robot successfully performed the diverse behaviors it had learned from the HiPHI data. It ran, sat down, crawled, carried a box, and pulled a suitcase, demonstrating that the high-precision recordings could be translated into actual physical action. This proved that the dataset does more than just describe movement; it provides a reliable foundation for teaching machines how to interact with the physical world. While the current dataset focuses on single-person movements and does not yet include interactions between multiple people or direct measurements of touch forces, it represents a significant step forward. By providing a large-scale, physically grounded record of human motion, HiPHI offers a new standard for developing robots that can move with the same grace and adaptability as the humans they are designed to assist.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →