PRISM: Precision and contact-rich Real-world Industrial Skill dataset with Multimodal sensing
This paper introduces PRISM, a large-scale, open-source multimodal dataset comprising over 5,000 teleoperated trajectories of 25 contact-rich industrial manipulation tasks, designed to bridge the gap between existing household-focused robotic datasets and the high-precision, force-regulated requirements of real-world manufacturing.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Robots have become remarkably adept at moving through the world, a capability largely built on the backs of massive libraries of video data. For years, researchers have taught machines to learn by watching thousands of examples of simple actions, like picking up a cup or stacking blocks in a kitchen or a living room. These everyday tasks are forgiving; if a robot misses a grip, it can usually try again without consequence. However, the industrial world operates under a much stricter set of rules. In a factory, assembling a machine part often requires the robot to push, slide, and twist with extreme precision, where a fraction of a millimeter of error can cause the entire process to jam. Success in these environments depends not just on seeing the object, but on feeling the subtle resistance of metal against metal, the slight give of a rubber seal, or the exact moment a gear clicks into place. For a long time, the data available to train robots simply did not capture these delicate, forceful interactions, leaving a gap between what robots can do in a lab and what they need to do in a real factory.
To bridge this divide, a team of researchers has introduced a new collection of data called PRISM, designed specifically to teach robots how to handle the messy, high-pressure reality of industrial assembly. The researchers gathered more than 5,000 detailed records of human operators performing complex tasks, such as plugging in delicate electronic components, sorting items on a moving conveyor belt, and packaging precision tools. Unlike previous datasets that focused on short, simple movements, these records capture long sequences of work where the robot must maintain constant contact with its environment. The data is rich with information, recording not only what the robot sees through multiple cameras but also the forces it exerts, the torque in its joints, and even the texture of the surfaces it touches through specialized sensors. By synchronizing all these different signals, the dataset provides a complete picture of what it feels like to perform a task that requires both a steady hand and a sensitive touch.
The team built this resource by equipping several different types of robots with advanced sensors and having eight volunteers perform the tasks using three distinct methods: wearing a robotic exoskeleton that mimics human movement, using handheld trackers, or controlling the robot through a virtual reality headset. This variety was intentional, allowing the researchers to see how different ways of guiding a robot affect the quality of the final data. They found that the method of control mattered significantly. When humans guided the robots using the exoskeleton, which provided a direct physical connection to the machine's joints, the resulting data led to robots that performed the tasks more successfully and with fewer errors. In contrast, data collected through virtual reality, where the operator viewed the scene through a flat screen, produced less reliable results, likely because the operator lacked the depth perception and physical feedback needed to guide the robot with the necessary precision.
To test the value of this new dataset, the researchers trained standard robot learning programs on the data and then sent them to a real robot to perform the same assembly tasks. The results showed that while the robots improved with more practice, they still struggled with the most difficult aspects of the work. When the robots were trained on a larger volume of data and given a broad overview of many different tasks before focusing on a specific one, they became more robust. They were better at recovering from small mistakes, such as a slight misalignment, rather than failing completely. However, the experiments also revealed a hard truth: even with this high-quality data, current learning methods still find it difficult to manage the split-second timing required for moving objects on a conveyor belt or to apply the exact amount of force needed to insert a part without breaking it. The robots could see the world and feel the contact, but they still lacked the deep, intuitive understanding of physics that a human worker develops over years of experience.
The work suggests that while large datasets are a powerful tool, they are not a magic solution for every industrial challenge. The PRISM dataset proves that capturing the full sensory experience of a task—sight, touch, and force—is essential for teaching robots to work in tight tolerances. It also highlights that the way we teach these machines matters just as much as the data itself; the physical connection between the human and the robot during training can make the difference between a robot that learns to adapt and one that simply copies movements it cannot truly understand. As the field moves forward, this dataset offers a realistic benchmark for measuring progress, showing that the path to truly versatile industrial robots will require not just more data, but better ways to interpret the complex, forceful dance of assembly.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.