RoboCousin: Build Your Own Simulation Playground for Robust Bimanual Robotic Manipulation
RoboCousin is an extensible simulation platform that automates the conversion of user-provided object observations into diverse, physics-accurate assets and expert trajectories, enabling scalable and robust data generation for bimanual robotic manipulation with effective sim-to-real transfer.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Teaching a robot to use two hands like a human is one of the most difficult challenges in modern robotics. To learn complex tasks like folding laundry, opening a jar, or carrying a heavy box, a robot needs to see thousands of examples of how to do it. In the real world, collecting these examples is slow, expensive, and often dangerous. If a robot drops a glass, it breaks; if it knocks over a stack of dishes, the cleanup is a chore. Because of this, scientists have turned to computer simulations. They build virtual worlds where robots can practice millions of times without breaking anything. However, a major problem has held this field back: most simulations are like closed rooms with fixed furniture. You can run the robot around the room as much as you want, but you cannot easily add a new object, like a specific coffee mug the user owns, or change the layout of the room. To do so, a human engineer usually has to spend hours manually rebuilding the object's shape, defining how heavy it is, and figuring out exactly where the robot's fingers should touch it. This manual work is so slow that it prevents the creation of the massive, diverse datasets needed to teach robots to be truly adaptable.
A team of researchers has introduced a new system called RoboCousin that solves this bottleneck by turning the simulation into an open, expandable workshop. Instead of forcing the robot to learn only from a pre-selected list of objects, this system allows a user to take a simple photograph of any object or a description of a room and instantly turn it into a realistic, interactive 3D model for the robot to practice with. The system does not just create a pretty picture; it automatically figures out the object's physical properties, such as its weight and how slippery its surface is, and it calculates the best places for the robot's grippers to grab it. This process transforms a flat image into a "digital cousin"—a virtual version that looks and acts like the real thing but is flexible enough to be mixed with other objects and backgrounds to create endless new training scenarios.
The researchers built this system on top of an existing simulation platform and added a pipeline that handles everything from the initial photo to the final robot movement. When a user provides an image of an object, the system first isolates it from the background and reconstructs its three-dimensional shape. It then uses artificial intelligence to guess the object's physical traits, such as its mass and friction, ensuring the robot interacts with it realistically. Crucially, the system automatically identifies "contact points," which are the specific spots on the object where a robot hand can safely and effectively grab it. In traditional setups, a human would have to manually mark these spots for every single object, a tedious process that limits how many objects can be added. In this new system, the computer does it instantly. The result is a library of over 3,000 different object instances and 50 different background environments, all ready for the robot to use.
To make the training even more robust, the system creates "digital cousins." While a "digital twin" is an exact, rigid copy of a real-world scene, a digital cousin is a variation that keeps the important relationships but changes the details. For example, if a robot needs to learn how to pick up a bottle from a table, the system can generate a scene where the bottle is a different color, the table is a different material, or the background is a kitchen instead of a living room, as long as the bottle is still sitting on the table in a way that makes sense. This approach teaches the robot the core concept of the task rather than just memorizing one specific setup. The system also scales up to the room level. It can plan a path for a mobile robot to navigate through a virtual house, approach a table, and then perform the manipulation task, all within a single continuous simulation. This allows the robot to learn not just how to move its arms, but how to move its whole body to get to the task.
The team tested whether this automated approach actually works in the real world. They generated over one million training trajectories using the system and then trained a physical robot with two arms to perform specific tasks. The results were striking. When the robot was trained on objects that the system had automatically reconstructed from images, it performed just as well, and in some cases better, than when it was trained on objects that had been carefully hand-crafted by human experts. For instance, in a task to pick up a bottle, the success rate jumped from 40 percent to 80 percent when using the new system's assets. Furthermore, when the robot was trained on a variety of "cousin" scenes rather than a single exact copy of a room, it became much better at handling new situations. In one test involving picking up an eyeglass case, a robot trained on a single exact scene failed completely, while the one trained on the varied cousin scenes succeeded 55 percent of the time.
The researchers also verified that the computer's automatic guesses about where to grab an object were accurate. They compared the system's suggested grab points against a set of points that had been carefully marked by humans. The computer's suggestions were slightly better, achieving a 76 percent success rate in simulation compared to 72 percent for the human-marked points. This proves that the system can replace the slow, manual process of annotating objects with a fast, automated one that is at least as reliable. By connecting the ability to turn real-world observations into simulation assets with the ability to generate diverse training scenarios, RoboCousin offers a practical path forward. It suggests that the future of robot learning does not require waiting for engineers to build every new object by hand, but rather relies on systems that can instantly understand and adapt to the messy, varied world we live in.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.