← Latest papers
🤖 AI

AI-based worker guidance in assembly and disassembly operations using multimodal ego/exo-centric data capture and structured task knowledge

This paper presents a data-centric approach that leverages multimodal egocentric and exocentric video and narration to extract structured task knowledge from expert demonstrations, enabling effective procedural documentation and context-aware guidance for assembly and disassembly operations.

Original authors: Vivek Chavan, Jörg Krüger

Published 2026-08-25
📖 5 min read🧠 Deep dive

Original authors: Vivek Chavan, Jörg Krüger

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the quiet hum of a factory floor, a skilled technician works on a complex machine. They know exactly which screw to loosen first, how to route a tangled cable without damaging it, and what to do when a part is missing or worn. This knowledge lives in their hands and their mind, built up over years of practice. For decades, engineers have tried to teach robots to do this work, but machines struggle with the messy, unpredictable reality of taking things apart. Unlike building a product from scratch, where every piece fits a perfect plan, taking an old device apart often means dealing with rust, broken clips, and parts that have shifted over time. Traditional automation, which relies on rigid, pre-programmed movements, often fails when the world does not follow the script. To solve this, researchers are turning to a different kind of intelligence: systems that can learn by watching and listening to human experts, capturing not just what they do, but why they do it.

A team of researchers at the Fraunhofer Institute and the Technical University of Berlin has developed a new way to capture this expert knowledge and turn it into a guide for workers. Their approach focuses on the difficult task of disassembling waste electrical and electronic equipment, such as old computers. Instead of trying to force a robot to figure out every step on its own, the system records a human expert performing the task while wearing smart glasses and speaking their thoughts aloud. The setup uses two cameras: one mounted on the expert's glasses to see exactly what they see, and a stationary camera in the room to show the whole workspace. As the expert works, they describe their actions, explain their reasoning, and point out potential mistakes. This creates a rich record that combines what the eye sees with the voice explains.

The researchers then use advanced computer programs to turn this recording into a structured map of the task. First, the system transcribes the spoken words into text, linking every sentence to the exact moment it was said. A powerful language model then reads this transcript to understand the logic of the job. It figures out which steps must happen before others, which steps can be done in any order, and which parts are independent. The result is a visual flowchart, or a precedence graph, that shows the correct sequence of actions. This map is not just a list of instructions; it understands the relationships between steps. For example, it knows that a cover must be removed before a specific screw can be reached, but that two different internal components can be taken out in any order.

Once this map is created and checked by a human to ensure it is correct, the system learns to recognize the actions in the video. It breaks the recording into small clips, each labeled with the specific step being performed. This allows the computer to understand the visual details of the work, such as how a hand holds a tool or how a component is pulled out. The system is then ready to assist a worker in real time. As a worker performs the task, the system watches their video feed, recognizes what they are doing, and checks their progress against the map. If the worker is on the right track, the system stays quiet. If they hesitate or make a mistake, the system can suggest the next valid step or show a short video clip of the expert performing that exact action correctly.

The team tested this method on the disassembly of a workstation computer, a task that involves many interacting parts and tools. They found that the system could successfully extract the logical structure of the job from just one expert demonstration. When they used recordings from four different experts, the system became even more reliable, better at recognizing actions and suggesting the correct next steps. In their tests, the system correctly identified the sequence of steps and the dependencies between them with high accuracy. It was able to guide a worker through the process, offering advice that fit the current situation. The results showed that even with a single example, the system could create a useful guide, and adding more examples made it robust enough to handle variations in how different people might perform the task.

This work suggests a path forward for industries that need to repair, recycle, or take apart complex products. By capturing the tacit knowledge of skilled workers and turning it into a digital guide, factories can support less experienced workers without needing to program every single robot movement. The system does not replace the human; instead, it acts as a knowledgeable partner that watches, understands, and offers help when needed. The researchers believe this approach could be expanded to help robots learn these tasks themselves, allowing machines to plan long sequences of actions based on human demonstrations. For now, the focus remains on helping human workers navigate the complexity of modern manufacturing, ensuring that valuable skills are not lost but are instead shared and preserved for the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →