← Latest papers
💻 computer science

Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning

This paper presents a zero-shot framework for long-horizon dexterous manipulation that leverages a vision-language model to ground language instructions into executable 3D task plans by fusing multi-view 2D keypoints into 3D space, enabling robust pick-and-place and tool-use execution on unseen objects without end-to-end training.

Original authors: Jisoo Kim, Sangwon Baik, Taeksoo Kim, Sungjoo Kim, Junyoung Lee, Mingi Choi, Hanbyul Joo

Published 2026-06-19
📖 5 min read🧠 Deep dive

Original authors: Jisoo Kim, Sangwon Baik, Taeksoo Kim, Sungjoo Kim, Junyoung Lee, Mingi Choi, Hanbyul Joo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a robot arm with a human-like hand that needs to clean up a messy room, but it has never seen this specific room before, and no one has taught it how to do this specific job. Most robots would need hours of training or thousands of examples to learn this. This paper presents a system that can do it zero-shot—meaning it figures it out on the fly, just by looking and listening, without any prior practice.

Here is how the system works, broken down into simple concepts:

1. The "Team of Photographers" vs. The "Single Snapshot"

Most robots look at a scene through one camera, like taking a single photo. If you try to guess where a cup is just from one photo, you might get the left-right position right, but you won't know exactly how far away it is. It's like trying to catch a ball in the dark with only one eye open; you might miss the depth.

This paper's system uses multiple cameras (like a team of photographers standing in different corners of the room).

  • The Problem: Each camera sees the object differently. One might see the handle of a spoon; another might see the bowl. One might be blocked by a book; another sees the whole thing.
  • The Solution: The system uses a "smart brain" (a Vision-Language Model) to ask every camera, "Where is the spoon?" and "Where should I grab it?"
  • The Magic Trick: It combines these different answers using two methods:
    1. Triangulation: Like how your two eyes work together to judge distance, it mathematically calculates the exact 3D spot where all the camera views agree.
    2. Ray Voting: If the cameras disagree (maybe one is blocked), the system shoots a "laser beam" from the main camera's view and asks the other cameras, "Did you see the object at this specific depth?" It picks the answer that gets the most votes. This ensures the robot doesn't grab thin air or miss the object because of a shadow.

2. The "Instruction Manual" vs. The "Muscle Memory"

Once the robot knows where the object is in 3D space, it needs to know how to move.

  • For Picking Things Up: The system breaks the big task ("Clean the room") into tiny, simple steps: "Grab the cup," "Move to the trash," "Drop it." It treats these like Lego blocks.
  • For Using Tools: This is the clever part. If the robot needs to use a broom or a kettle, it doesn't need to learn how to sweep from scratch. It has a "Bag of Atomic Actions" in its memory. Think of this as a library of pre-recorded dance moves for tools.
    • If the instruction is "Sweep the floor," the robot finds the "Sweeping" dance move in its library.
    • It then aligns that dance move to the current room. If the broom is on the left and the trash is on the right, it stretches and rotates the pre-recorded "sweeping" motion to fit that specific geometry.

3. The "Safety Check" and "Do-Over"

Robots often fail because they try to grab something and hit a wall, or they try to put a pot on a stove that is too high.

  • Affordance Regions: Instead of just grabbing a single point, the system figures out the "safe zone" on an object. For a handle, it knows the whole handle is a good place to grab, not just one pixel.
  • Collision Avoidance: Before moving, it simulates the path to make sure the robot hand won't crash into the table or the wall.
  • Closed-Loop Verification: This is the robot's "reality check." If the robot tries to pick up a cup and the camera sees it didn't move, the system doesn't just keep trying the same failed move. It stops, re-evaluates the scene, and re-plans. It's like a human saying, "Oops, I missed, let me try a different angle," rather than blindly repeating the mistake.

The Results

The authors tested this on a real robot with a dexterous hand in a real room.

  • Better Accuracy: It was much better at finding the exact 3D location of objects than robots that only look at one camera.
  • Zero-Shot Success: While other advanced robots (trained on 30 specific examples) failed completely at new tasks, this system succeeded at tasks like "throw away trash," "place a pot on a stove," and "sweep with a broom" without ever being trained on those specific tasks before.
  • Long-Horizon Tasks: It could string together multiple steps (like organizing a whole table) and recover if it made a mistake in the middle.

In Summary

This paper describes a robot that acts like a smart, adaptable human. Instead of memorizing every possible way to move, it uses a smart brain to understand language and 3D space from multiple angles, pulls pre-made "tool dances" from a library, and constantly checks its own work to fix mistakes on the fly. It proves you don't need to train a robot for every single job if you give it the right tools to reason about the world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →