← Latest papers
💻 computer science

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

This paper introduces ReactHuman, the first physics-grounded benchmark that evaluates embodied multimodal LLMs on their ability to make immediate, safety-critical reactive decisions in simulated household environments, revealing that current models struggle with intuitive physics and safety despite scaling.

Original authors: Yizhan Li, Jianxin You, Mengyang Xiong, Yinhuan Chen, Zicheng Zhao, Dekun Wu, Dongqing Zhang, Bang Liu

Published 2026-09-11
📖 4 min read☕ Coffee break read

Original authors: Yizhan Li, Jianxin You, Mengyang Xiong, Yinhuan Chen, Zicheng Zhao, Dekun Wu, Dongqing Zhang, Bang Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a kitchen where a heavy pan slips off a counter, or a hallway where a tall wardrobe begins to topple. In a fraction of a second, a human brain processes the visual danger, predicts where the object will land, and commands the body to either catch it or step aside. This split-second reaction is not just a reflex; it is a complex fusion of understanding what an object is, knowing how heavy it is, and calculating exactly where it will go. As scientists work to build robots that can live and work alongside us in our homes, they face a critical question: can artificial intelligence learn to react with this same speed and safety? Current tests for these intelligent systems often ask them to answer questions about videos or plan long-term tasks like cleaning a room, but these methods do not test whether a machine can make a life-saving decision the moment a hazard appears.

A new study introduces a rigorous test called ReactHuman, designed to see if artificial intelligence can act like a competent human when faced with sudden physical dangers. The researchers built a simulated world where a digital robot stands in a room and watches a series of dangerous events unfold, such as a falling knife, a sliding shelf, or a ball bouncing toward it. The system pauses the simulation just before the disaster happens, showing the robot a few seconds of the event from multiple camera angles. The artificial intelligence, acting as the robot's brain, must then decide what to do and issue a specific command: walk to a new spot, reach out to grab the object, or stay still. Unlike previous tests where a computer might simply choose the right answer from a list, this system forces the robot to actually perform the action in a physics simulation. If the robot decides to catch a falling object, the simulation runs to see if its hand actually meets the object in time. If it decides to dodge, the simulation checks if the robot successfully moves out of the way.

The researchers created over one thousand unique scenes covering seventeen different types of sudden events, ranging from simple drops to complex chain reactions where one falling item knocks into another. They even included "trick" objects that look like one thing but behave like another, such as a foam anvil that looks heavy but is light, or a steel apple that looks like fruit but is heavy and dangerous. This setup tests whether the robot relies on how things look or on how they actually move. The team evaluated seven different advanced artificial intelligence models on this task. The results were sobering: the models failed to react correctly about one time in three. When the correct action was to move away from danger, the robots often froze in place or tried to catch the object, leading to unsafe outcomes.

Perhaps most revealing was that the robots did not seem to learn from their mistakes as they grew larger or more complex. The biggest, most powerful models were not significantly safer or more accurate than the smaller ones. Instead of analyzing the specific scene in front of them, each model seemed to have a fixed personality: some always preferred to run away, while others always tried to grab things, regardless of whether that was the right move. The robots also struggled to judge speed. They could tell if something was moving, but they could not tell if it was moving slowly or quickly, often treating a slow-moving hazard as if it were not a threat at all. When the robots did choose the right action, they frequently failed to execute it correctly, reaching for an object but stopping their hand a meter away from where it needed to be.

The study concludes that while these artificial intelligence systems are getting better at understanding the world, they are still far from being able to react safely to sudden physical hazards. The research highlights that simply making models larger does not solve the problem of safety. To build robots that can truly live in our homes, developers will need to teach them to trust what they see happening in the moment rather than relying on what an object looks like, and to learn how to judge speed and distance with the same instinctive precision that humans possess. Until then, the gap between a machine that can describe a falling object and one that can safely catch it remains wide.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →