CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
CoToGrasp is a novel, object-agnostic generative framework that synthesizes diverse, stable dexterous grasps conditioned on specific contact topologies by learning in a feature-based canonical workspace, thereby achieving zero-shot generalization to unseen objects without requiring expensive annotated datasets.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Robots have long been masters of moving heavy objects, but they struggle with the delicate art of holding things the way a human does. For decades, engineers treated grasping as a simple geometry problem: find a spot where a robot hand can hold an object without dropping it. This approach works well for picking up a box to move it from one place to another, but it fails when the robot needs to perform a specific task, like turning a screwdriver or holding a cup by its handle. In these situations, how the object is held matters just as much as that it is held. A robot needs to know not just where to grab, but which fingers to use and how to arrange them to support a future action. This shift from simple stability to functional intent is the next great challenge for robots that are meant to live and work alongside people.
A team of researchers has developed a new system called CoToGrasp to solve this problem. Instead of teaching a robot to memorize how to hold thousands of different objects, they taught it to understand the "shape" of a grip itself. The researchers realized that while every object in the world is different, the way a hand touches an object follows a limited set of patterns. They borrowed a classification system originally designed for human hands, which sorts grasps into categories like "precision" (using just the fingertips for fine control), "power" (wrapping the whole hand around something for strength), and "object-specific" (using a unique grip for a specific tool). By focusing on these contact patterns rather than the objects themselves, the system can learn to generate the correct grip for a tool it has never seen before.
The core of this breakthrough is a method that separates the idea of a grip from the object being held. Imagine a robot hand that learns to recognize the feeling of a pinch or a squeeze in its own frame of reference, completely independent of what it is holding. The researchers created a virtual workspace anchored to the robot's hand, where they mapped out exactly which parts of the fingers should touch an object for a specific task. When the system needs to grasp a new item, it projects the object's surface features into this hand-centered space. It then asks, "If I want a precision pinch, where on my fingers should the contact happen?" The system generates a map of contact points based on the desired task, and then the robot figures out how to move its joints to match that map. This allows the robot to ignore the confusing details of the object's shape and focus entirely on the functional goal.
To test this idea, the researchers trained their system using only data about the robot hand itself, never showing it a single picture of an object during the learning phase. This "object-agnostic" approach meant the system did not need to memorize millions of object-hand pairs, which is usually required for such tasks. Instead, it learned the intrinsic capabilities of the hand. When they tested the system on a large dataset of unseen objects, it successfully synthesized grasps that matched specific functional categories. The results showed that the system could produce a wide variety of grips, from delicate precision holds to strong power grips, without defaulting to the same few easy solutions that other robots often use. In fact, the system generated grasps with a much higher variety of functional types than previous methods, which tended to get stuck in a loop of only using simple, enveloping grips.
The researchers also compared their system to other advanced planners that try to follow similar rules. While those other systems often failed when the object's shape didn't perfectly match a pre-programmed template, CoToGrasp adapted smoothly. It successfully generated valid grasps for difficult, highly constrained tasks, such as holding a tool in a very specific way, with a success rate that outperformed the competition. The system was able to filter out impossible grasps before the robot even tried to move, ensuring that only physically viable options were selected. This filtering process is crucial because it prevents the robot from wasting time trying to grab something in a way that is physically impossible.
Finally, the team took the system out of the computer simulation and onto a real robot arm equipped with a multi-fingered hand. They placed various everyday objects on a table and asked the robot to grasp them using specific contact patterns. The robot successfully executed these complex grips in the real world, holding objects steady and demonstrating the physical viability of the generated plans. The experiments confirmed that the system could translate a high-level instruction like "use a precision grip" into a concrete, stable physical action on an object it had never encountered before. This work suggests that by teaching robots to understand the functional logic of a grip rather than just the geometry of an object, we can build machines that are far more capable of the nuanced, task-oriented manipulation required in our daily lives.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.