← Latest papers
💻 computer science

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval

This paper introduces MRBench, a comprehensive human motion-text retrieval benchmark featuring diverse, balanced, and multi-granular data to address limitations in existing datasets, alongside a lightweight granularity-aware model that significantly improves cross-domain generalization and retrieval performance across different description levels.

Original authors: Fulong Liu, Liang Xu, Chengqun Yang, Yuhao Zhang, Yichao Yan, Xiaokang Yang

Published 2026-08-11
📖 3 min read☕ Coffee break read

Original authors: Fulong Liu, Liang Xu, Chengqun Yang, Yuhao Zhang, Yichao Yan, Xiaokang Yang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to understand human movement, not just by watching it, but by listening to how we describe it. This is the world of human motion-text retrieval, a field where computers try to match a video of a person dancing or walking with the specific words used to describe that action. Think of it like a super-powered librarian who doesn't just look at the book's title, but understands the story inside to find the exact match for a request like, "Show me the video where the person trips over their own shoelaces." For a long time, scientists have been building libraries of these motion videos and descriptions to train their robots. However, there's a catch: if the library only contains videos of people walking in a straight line in a white room, and the descriptions are all just "person walks," the robot learns a very narrow, boring version of the world. It might be great at finding that one specific walk, but it will be completely lost if you ask it to find someone doing a backflip in a park or a clumsy stumble in a kitchen.

This is exactly the problem a team of researchers from Shanghai Jiao Tong University and the Beijing Zhongguancun Academy decided to tackle. They realized that the old libraries of motion data were too simple, too repetitive, and didn't reflect the messy, diverse reality of how humans actually move. So, they built a brand new, massive library called MRBench. Instead of just "walking," their library includes 3,390 different movements ranging from indoor dance moves to wild, real-world actions captured from videos and even computer-generated animations. They covered 118 different categories of movement, ensuring that no single type of motion dominated the collection. But they didn't stop at just the videos; they also rewrote the descriptions. They created three levels of detail for every single motion: a short, punchy summary (like "person falls"), a standard description (like "person falls backward to the floor"), and a super-detailed, fine-grained account (like "person starts standing, leans back, collapses, and lands motionless"). This allows them to test if a computer can understand a simple command versus a complex, detailed story.

When they put their new library to the test, they found that the old, popular computer models were in for a rude awakening. These models, which seemed smart when tested on the old, simple libraries, struggled significantly when faced with MRBench's diversity. They were like a student who memorized the answers to a specific practice test but failed when the questions were phrased differently or asked about new topics. The researchers discovered that these models were very sensitive to how detailed the text was; they often got confused when the description wasn't exactly what they were used to. To fix this, the team designed a new, lightweight model that acts like a flexible translator. It keeps a solid foundation for understanding standard descriptions but adds special "adapters" that can quickly adjust to understand short summaries or long, detailed stories without losing its way. By testing this new approach, they showed that it's possible to make computers better at understanding the full spectrum of human movement and language, from a quick "jump" to a detailed play-by-play of a gymnastics routine, without forgetting how to handle the basics.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →