← Latest papers
💻 computer science

HumanCLAW: Can Vision-Language Models Act Through a Body?

This paper introduces HumanCLAW, a framework that decouples vision-language model decision-making from motor execution to evaluate embodied intelligence, revealing that current models fail at long-horizon physical tasks primarily due to a lack of embodied self-awareness rather than recognition capabilities.

Original authors: Siyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li
Published 2026-07-30
📖 4 min read☕ Coffee break read

Original authors: Siyao Li, Jiawei Gu, Shuai Liu, Kairui Hu, Zekun Li, Linjie Li, Chengcheng Tang, Po-Chen Wu, Ivan Shugurov, Lingni Ma, Michael Zollhoefer, Sizhe An, Abhay Mittal, Amy Zhao, Ranjay Krishna, Manling Li, Ziwei Liu, Chuan Guo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a brilliant, super-smart robot to play a video game. But this isn't just any game; it's a game where the robot has to control a real, physical body that can trip, fall, bump into walls, and knock over vases. In the world of artificial intelligence, we have two main types of "brains" for robots. One type is like a pilot who has memorized millions of specific flight paths; it's great at doing exactly what it's seen before, but if the wind blows differently, it might crash. The other type is a "Vision-Language Model" (VLM). Think of this as a robot with a massive library of books and a camera for eyes. It can read a story, look at a picture, and understand complex ideas like "find the red sofa and sit on it." It's amazing at reasoning and talking, but it has never actually moved a body before.

The big question scientists are asking right now is: Can this book-smart robot actually do things in the real world? Can it take a high-level idea like "go to the kitchen" and turn it into the tiny, split-second muscle movements needed to walk without falling over? Usually, when a robot fails, it's hard to tell if the "brain" made a bad decision (like walking into a wall) or if the "body" just couldn't do the job (like losing balance). To solve this mystery, researchers needed a way to separate the thinking from the moving, so they could see exactly where the robot's brain was getting stuck.

This is where a new study called HumanCLAW comes in. The researchers built a special testing ground to see if these smart AI brains could control a human-like body in a physical world. They didn't train the AI to move; they just gave it a camera, a goal, and a list of basic moves like "walk forward," "turn," or "sit." Then, they let the AI try to navigate through 41 different virtual houses to find objects and sit on them.

The results were a bit of a shocker. Even the smartest, most advanced AI models out there today failed almost every time. The best model only succeeded in about 16.8% of the attempts. That means it failed more than 83% of the time.

Here is the weird part: The AI wasn't failing because it was "blind." When the target object (like a sofa) actually appeared on the screen, the AI almost always recognized it. It knew what it was looking at. The problem wasn't seeing; it was knowing where it was.

The researchers discovered that these AI models suffer from a severe case of "body amnesia." They are like a ghost that can see the furniture perfectly but has no idea where its own legs are. When the AI tried to walk, it would often keep walking even after it had already arrived at the sofa, or it would stop way too early, thinking it had arrived when it was still far away. When it tried to sit, it would often just lower its body into thin air, sitting on nothing because it didn't realize the sofa wasn't directly underneath it. It would also walk right into walls or trip over obstacles without ever realizing it had hit them.

The study shows that while these AI models are incredible at describing the world, they are terrible at understanding their own place in it. They can tell you what a chair looks like, but they can't figure out how to sit on it without falling over. The researchers conclude that to make robots that can truly act in the real world, we need to teach them not just to see, but to "feel" where their body is and how it interacts with the world around them. Until we solve this "embodied self-awareness" problem, even the smartest AI brains will remain clumsy ghosts in a physical machine.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →