ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning
ADEPT is a large-scale reinforcement learning framework that accelerates the development of human-level dexterity in high-degree-of-freedom robots by pre-training on generic tasks and employing a stable post-training recipe to enable zero-shot sim-to-real transfer for solving complex, long-horizon manipulation tasks from raw visuo-tactile perception.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Robots have long been masters of the factory floor, moving with rigid precision to assemble cars or stack boxes. Yet, when asked to perform the fluid, delicate tasks of daily life—like picking up a fragile egg, turning it over, and placing it gently into a bowl—most machines falter. This is the challenge of dexterity. It requires a robot to coordinate dozens of moving parts simultaneously, reacting to touch and sight in real time to manage the chaotic physics of objects slipping, rolling, or colliding. For years, researchers have tried to teach robots these skills by programming them with rigid rules or showing them human demonstrations, but these methods often fail when the environment changes even slightly. A new approach, however, suggests that the secret to robotic dexterity lies not in teaching a robot every specific task from scratch, but in giving it a broad, foundational experience with the physical world first.
A team of researchers has developed a new framework called ADEPT that teaches robots to learn complex manipulation skills by first mastering a generic game of "repositioning" objects. Imagine a robot hand that spends months in a virtual world simply picking up random shapes—cylinders, cubes, spheres—and moving them to different spots on a table. It learns to reach, grasp, lift, and turn these objects without any specific goal other than moving them successfully. This phase builds a deep, intuitive understanding of how objects behave and how the robot's own fingers should move to control them. Once this foundation is laid, the researchers introduce a specific new task, such as inserting a peg into a hole or placing a plate into a dish rack. Instead of starting over, the robot uses its prior experience as a guide. The researchers found that this method allows the robot to skip the initial, clumsy learning phase and immediately apply its general skills to the new, specific challenge.
The power of this approach becomes clear when the robot faces a task it has never seen before. In their experiments, the researchers tested the system on two different robot arms equipped with multi-fingered hands: one with twenty-three moving joints and another with twenty-nine. They trained the robots in a high-fidelity simulation where they could run thousands of trials simultaneously. The robots first learned to reposition a variety of simple shapes. When the researchers then asked the robots to perform a difficult task—inserting a peg into a board, a challenge that requires precise alignment and handling of contact forces—the robots did not start from zero. They immediately recognized the need to reach, grasp, and lift the peg, skills they had already mastered. The only part they had to learn was the final, delicate insertion.
However, simply taking a robot trained on one task and asking it to do another often causes it to forget what it learned. The researchers discovered that if they tried to fine-tune the robot's behavior directly, the new instructions would overwrite the old skills, causing the robot to lose its ability to grasp or lift properly. To solve this, they developed a careful training recipe. First, they used a technique to gently nudge the robot's new behavior to look like the old, successful behavior, ensuring the core skills remained intact. Then, they slowly adjusted the robot's understanding of what a "good" move looks like for the new task, allowing it to adapt without losing its balance. Finally, they trained a second, simpler version of the robot that could rely only on what it saw through cameras and what it felt through its fingertips, stripping away the internal data the robot used during training. This "student" robot was then sent to the real world.
The results were striking. In the real world, these robots could successfully pick up a peg, reorient it in their hand, and insert it into a hole, all while relying only on visual and tactile feedback. They could handle a symmetric star-shaped peg and a much more difficult, asymmetric square-and-round peg, which requires a single, precise orientation to fit. The system also learned to grasp a flat plate, flip it over, and place it into a dish rack, a sequence of actions that the initial training had never explicitly shown. The robots performed these tasks in five to ten seconds, a speed comparable to a human, whereas previous methods using simpler grippers and external fixtures took twenty to seventy seconds.
Crucially, the researchers found that the robots developed natural, human-like ways of holding objects. Without being told how to hold the peg, the robots learned to wrap their fingers around it in a stable, multi-point grip, much like a human hand. Robots trained without the initial repositioning phase, by contrast, often developed awkward, unstable ways of holding the object that worked for a moment but failed under pressure. The study suggests that by giving a robot a broad, general experience with the physics of objects first, it can learn new, specific skills much faster and more reliably than by trying to learn everything at once. This approach does not just make robots faster; it makes them more robust, capable of handling the unpredictable nature of the real world with a level of dexterity that was previously out of reach.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.