← Latest papers
💻 computer science

ShapeGrasp: Simultaneous Visuo-Haptic Shape Completion and Grasping for Improved Robot Manipulation

ShapeGrasp is a novel robotic system that iteratively combines visual estimation with tactile feedback during grasping to simultaneously refine 3D object shape reconstruction and improve grasp success rates on unfamiliar objects.

Original authors: Lukas Rustler, Matej Hoffmann

Published 2026-07-01
📖 4 min read☕ Coffee break read

Original authors: Lukas Rustler, Matej Hoffmann

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a robot trying to pick up a mysterious object it has never seen before, like a weirdly shaped tool or a crumpled box. If the robot only uses its eyes (a camera), it sees a flat, incomplete picture. It might guess the back of the object is smooth, but in reality, there's a handle sticking out. If the robot tries to grab it based on that guess, it might miss, drop the object, or crush it.

ShapeGrasp is a new method that teaches robots to learn by doing, much like how a human would. Instead of just looking and guessing, the robot "feels" its way to a solution.

Here is how it works, broken down into simple steps:

1. The "Blindfolded" Start

The robot starts with a single snapshot from a camera. It uses a smart computer program to guess what the rest of the object looks like. Think of this like looking at a puzzle piece and trying to imagine the whole picture. The robot creates a 3D model, but it's just a guess.

2. The "Trial and Error" Grab

The robot tries to grab the object based on that guess.

  • If it succeeds: Great! It lifts the object.
  • If it fails: The object might slip, or the robot might realize, "Wait, my guess was wrong; the object is heavier or shaped differently than I thought."

3. The "Sensory Update"

This is the magic part. Every time the robot tries to grab the object, it learns two new things:

  • Touch: If the robot's fingers actually touch the object, its sensors tell it, "Here is the exact surface of the object." It's like the robot running its fingers over a statue to feel the bumps.
  • Empty Space: Even if the robot misses, its fingers take up space. The robot knows, "My fingers couldn't go here because the object must be in the way." It uses this "empty space" data to refine its mental map.

4. The "Do-Over"

If the first grab failed, the robot doesn't just try the same thing again. It takes all the new touch and space information, updates its 3D model of the object to make it more accurate, and then tries a new grab based on this better map. It keeps doing this loop—grab, feel, update, re-grab—until it succeeds.

Why is this special?

Most robots either try to build a perfect 3D model before they touch anything, or they just grab blindly. ShapeGrasp is the first to update its 3D model in the real world after a physical grab.

Think of it like a sculptor. A traditional robot is like someone who tries to carve a statue based on a blurry photo and hopes for the best. ShapeGrasp is like a sculptor who chisels a little, steps back, feels the stone, realizes they chiseled too much, and then adjusts their plan before chiseling again.

The Results

The researchers tested this on two different robot arms with different types of "hands" (grippers):

  • A two-fingered hand (like a pincer).
  • A three-fingered hand (more like a human hand).

They found that ShapeGrasp was much better at picking up objects than previous methods.

  • The two-fingered hand succeeded 91% of the time.
  • The three-fingered hand succeeded 84% of the time.

Crucially, every time the robot tried to grab the object, the 3D model of that object became more accurate and detailed. By the end, the robot knew exactly what the object looked like, which helps it handle the object carefully later on (for example, if it needs to hand the object to a human).

The Trade-off

The paper notes that this process takes a bit of time (about 5 to 9 seconds per object) because the robot has to think, simulate, and move carefully. However, the trade-off is worth it because it is much more reliable than guessing.

In short, ShapeGrasp turns a robot from a "guess-and-hope" machine into a "learn-and-adapt" machine, using its own mistakes to build a better understanding of the world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →