IMBench: A Benchmark for Intuitive Robotic Manipulation
This paper introduces IMBench, a comprehensive benchmark comprising 35 tasks and 14K trajectories designed to evaluate the integrated capability of "intuitive manipulation"—combining physical reasoning, perception, and action generation—revealing significant gaps in current vision-language and vision-language-action models' ability to plan and execute complex, constraint-rich robotic tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where robots are like incredibly talented actors who have memorized millions of scripts. They can recite lines perfectly and move their arms with precision if the stage is set exactly as they've seen before. But ask them to improvise a scene where the floor is slippery, or a prop is heavier than it looks, and they often freeze or stumble. This is the current state of robotics: they are great at following instructions but struggle with "intuition."
To understand why this matters, we need to look at two superpowers humans have that robots currently lack. The first is intuitive physics. This isn't a complex math class; it's the brain's ability to guess how things will move, fall, or bounce just by looking at them. It's knowing that a tall, wobbly stack of books will tip over if you push it too hard, or that a thin plate on a table can't be grabbed from the top without sliding it to the edge first. The second is means-end reasoning, which is the ability to figure out how to get from where you are to where you want to be, even if the direct path is blocked. It's the difference between trying to grab a cookie from a high shelf and realizing you need a chair first.
Scientists care about this because for robots to truly help us in messy, real-world homes, they can't just be pre-programmed machines. They need to understand the physical world the way we do, so they can adapt when things don't go according to plan. If a robot can't figure out that a cup is too slippery to hold, or that a door is stuck, it's not very useful.
Enter IMBench, a new "test drive" for robots designed by researchers at Manav Robotics. Think of IMBench not as a standard exam with multiple-choice questions, but as a series of tricky, real-life puzzles. The researchers built 35 different scenarios—like balancing a rod on a tiny ridge, using a hook to pull a toy out of reach, or stacking weirdly shaped blocks without them toppling over. These tasks are designed to trick robots that rely only on memorizing patterns. They force the robot to pause, think about the physics, and come up with a new plan on the fly.
The paper tests three different types of "brains" on these puzzles. First, they asked powerful AI chatbots (Vision-Language Models) to just look at the scene and explain what's happening. Surprisingly, these AI models were pretty good at spotting the rules. They could tell you, "Hey, that plate is too thin to grab from the top," about 74% of the time. They understood the theory.
However, the plot thickens when the researchers asked these same AIs to make a plan to solve the problem. The success rate dropped. The AI could explain the problem but struggled to figure out the exact steps to fix it. It was like a student who could explain the rules of soccer perfectly but couldn't figure out how to kick the ball into the goal.
Finally, the researchers tested actual robot control systems (policies) that were supposed to execute the actions. Here, the results were quite humbling. Even the most advanced robot policies, which had been trained on thousands of examples, failed miserably on these intuitive tasks. When asked to balance a rod or use a tool, they succeeded less than 25% of the time, and often not at all. The paper suggests that while these robots are getting better at copying human movements, they are still missing that crucial spark of physical intuition. They can mimic the action, but they don't truly understand why the action works.
The researchers conclude that we are missing a key piece of the puzzle. Current robots are like actors who can memorize a script but can't improvise. IMBench shows us exactly where they are failing: in the gap between understanding a physical problem and actually solving it with their hands. It's a reminder that for robots to become our true partners, they need to learn not just how to move, but how to think about the world they are moving in.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.