← Latest papers
💻 computer science

LabDex: A Hierarchical Benchmark for Dexterous Manipulation in Laboratories

This paper introduces LabDex, a large-scale, hierarchical benchmark and dataset that unifies real-world and simulation environments to systematically evaluate and train robotic policies for dexterous manipulation across atomic skills, compositional tasks, and long-horizon experiments in chemistry laboratories.

Original authors: Zhipeng Tang, Sihang Chen, Sha Zhang, Peihao Yang, Yan Liu, Wentao Zhao, Xinrui Liu, Rui Huang, Wensheng Du, Yuting Huang, Jiajun Deng, Lidian Wang, Yuan Zhang, Yanyong Zhang

Published 2026-08-20
📖 6 min read🧠 Deep dive

Original authors: Zhipeng Tang, Sihang Chen, Sha Zhang, Peihao Yang, Yan Liu, Wentao Zhao, Xinrui Liu, Rui Huang, Wensheng Du, Yuting Huang, Jiajun Deng, Lidian Wang, Yuan Zhang, Yanyong Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where the most delicate, repetitive, and precise work in a chemistry lab is done not by a human hand, but by a machine. For decades, scientists have dreamed of automating these spaces to speed up discovery, but the robots we have built so far are often clumsy. They are equipped with simple, pincer-like grippers that can grab a box or a block, but they struggle with the nuanced, fluid movements required to handle a glass test tube, pour a liquid without spilling, or stir a solution with a thin rod. To move forward, researchers needed a way to teach machines the kind of dexterity a human possesses—the ability to use fingers to feel, adjust, and manipulate objects with care. This is the challenge of "dexterous manipulation," a field that seeks to give robots hands that can do more than just hold things; they must be able to perform complex, multi-step tasks that require a gentle touch and a steady hand.

In a new study, a team of researchers has created a massive new testing ground to solve this problem, specifically for chemistry laboratories. They call it LabDex. Instead of just asking a robot to pick up a single object, the researchers organized the entire scope of laboratory work into a clear, three-level hierarchy. At the bottom are "atomic skills," which are the tiny, fundamental moves like grasping a glass rod, pouring water, or shaking a tube. The next level up is "compositional skills," where the robot must string several of those tiny moves together to complete a single function, such as picking up a beaker, pouring its contents, and setting it down. Finally, at the top are "long-horizon workflows," which are complete experiments that require the robot to perform a long sequence of different actions to achieve a final scientific result. By breaking the work down this way, the team could see exactly where a robot succeeds and where it fails, rather than just looking at whether the final experiment worked or not.

To build this benchmark, the researchers constructed a unified system that works both in the real world and in a computer simulation. In their physical lab, they used a robot arm equipped with a sophisticated, multi-fingered hand that looks and moves much like a human hand. A human operator, wearing special data gloves and using motion-tracking devices, teleoperated the robot to perform these tasks by physically guiding it through the motions. The robot recorded every movement, every camera angle, and every joint position, creating a library of thousands of demonstrations. They then translated these real-world lessons into a virtual environment, allowing them to test the robot's learning in a safe, repeatable digital space before trying it again on the actual hardware. This dual approach ensures that the data is not just a collection of numbers, but a faithful representation of the physical realities of a chemistry lab.

The researchers then put several of the most advanced robot learning systems to the test using this new benchmark. They watched as these systems tried to learn the tasks from the demonstration data. The results were revealing. While some of the newer, more advanced models could successfully perform the simple, single-step tasks—like picking up a beaker or placing it on a table with high accuracy—they struggled significantly when asked to chain those steps together. When the tasks became more complex, requiring the robot to pour liquid from one container to another and then return the container, the success rates dropped sharply. The most capable model managed to complete an average of 0.73 atomic skills per multi-part task, but none of the models was able to finish the entire sequence of a full experiment on its own.

A closer look at the failures showed that the problem was not just about making a mistake once; it was about how a small error early in the process ruined everything that came after. For instance, if a robot managed to pick up a funnel but held it at a slightly wrong angle, it would fail to insert it into a flask, and the entire experiment would stop. The study found that the robots were particularly bad at handling objects that were long, thin, or made of glass, such as test tubes and graduated cylinders. These objects are sensitive to exactly where the fingers touch them, and the robots lacked the fine control to adjust their grip in real time. Furthermore, when the researchers added extra objects to the table that looked similar to the target, the robots became confused, often grabbing the wrong item or failing to act altogether.

The team also tested how well these robots could learn from more data. They found that simply showing the robot more examples of how to do a task helped, but only up to a point. The most advanced models improved significantly when they saw hundreds of demonstrations instead of just a few, suggesting that these systems need a vast amount of experience to learn the subtle differences between a successful pour and a spilled one. However, even with this extra data, the robots still could not reliably complete the full, long experiments that a human chemist could do with ease. The study concludes that while we have made progress in teaching robots to perform individual, simple moves, the ability to combine those moves into a stable, reliable, and long-lasting workflow remains a major hurdle.

This work does not claim to have solved the problem of the autonomous laboratory. Instead, it provides a clear map of where the road is blocked. By showing exactly which skills are missing and how errors in one step cascade to ruin the whole process, the researchers have given the scientific community a precise target for future improvements. The LabDex benchmark serves as a standard yardstick, allowing different teams to measure their progress against the same set of difficult, real-world tasks. It is a reminder that building a robot that can truly work in a lab requires more than just a strong arm; it requires a hand that can think, feel, and adapt, just like the human hands that have done this work for centuries.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →