MS-MEM: Multi-Skill Manipulation-Enhanced Mapping via Uncertainty- and Disturbance-Aware Action Selection
This paper proposes MS-MEM, an evidential framework that integrates active viewpoint selection, object pushing, and grasping with a collateral disturbance constraint to enhance mapping accuracy in cluttered environments while minimizing unnecessary scene changes.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Robots that work in our homes and warehouses face a problem that is deceptively simple: they cannot see everything. In a cluttered room, objects hide behind one another, creating blind spots that confuse a robot's sensors. To find a specific item, a robot must do more than just look; it must interact with the world. It might push a box aside to see what is behind it, or move to a new angle to get a better view. This field of study, known as active perception, relies on the idea that a robot learns best by touching and moving things, not just by staring at them. However, there is a catch. Every time a robot pushes an object, it changes the scene. If it pushes too hard or too often, it might scatter a carefully arranged shelf, making it harder to understand the layout later. The challenge for engineers is to teach a robot how to be curious enough to find hidden things, but careful enough not to make a mess in the process.
Researchers at the Karlsruhe Institute of Technology, the University of Bonn, and the Technical University of Darmstadt have developed a new system to solve this balancing act. They call it MS-MEM, a framework designed to help robots map out tight, cluttered spaces like shelves while keeping the scene as undisturbed as possible. The system allows a robot to choose between three different types of actions: moving to a new viewpoint to see more, pushing objects to reveal what is hidden, or picking up and removing an object that is blocking the view. The core innovation is that the robot does not just pick the action that reveals the most information; it also calculates the cost of that action. It asks itself, "If I push this box, how much will I disturb the rest of the shelf?" and "Is the new information worth the mess I might make?"
The team built their system on a foundation of uncertainty. Instead of treating the map of the room as a fixed picture, the robot maintains a "belief" about what is where, complete with a measure of how unsure it is. When the robot is very confident about an area, it knows not to touch it. When it is unsure, it knows it needs to investigate. To make these decisions, the researchers introduced a new way for the robot to understand grasping. Previous systems often struggled to predict how a robot's gripper would interact with an object when the view was partial or the object was in a tight spot. This new system teaches the robot to estimate not just where to grab, but how confident it is in that grab, and how the orientation of the object might be uncertain. This allows the robot to fuse information from multiple angles, refining its understanding of a grasp over time as it moves around the object.
In their experiments, the researchers tested this system in a simulated environment that mimicked a real-world shelf filled with randomly placed objects. They compared their new multi-skill approach against older methods that relied on only one type of action, such as only pushing or only grabbing. The results showed that combining these skills was far more effective. By using pushing to separate objects and create space, and then using grasping to remove specific blockers, the robot was able to build a more accurate map of the scene. However, the most significant finding came from the system's ability to limit disturbance. When the researchers added a specific constraint to the robot's decision-making process—a rule that penalized unnecessary changes to the scene—the robot became much more careful. It still found the hidden objects, but it moved far fewer items than the systems that did not have this rule. In simulations, the new system reduced the total distance objects were moved by a significant margin compared to methods that ignored scene disturbance, while still achieving high accuracy in mapping the shelf.
The team also tested the system in the real world using a physical robot arm equipped with a camera and a gripper. They placed the robot in front of a shelf containing 69 objects across five different challenging setups. The robot had to find and identify as many objects as possible. The system that combined pushing, grasping, and careful movement found more objects than the systems that used only one skill. It successfully identified 44 objects, compared to 40 for the pushing-only system and 38 for the grasping-only system. Crucially, it did this while causing less physical displacement of the items on the shelf. The grasping-only system moved the least, but it failed to find many hidden items because it was too conservative. The pushing-only system found many items but scattered the shelf significantly. The new system struck the best balance, finding the most objects while keeping the scene relatively orderly.
This work demonstrates that for robots to operate effectively in human spaces, they must be able to reason about the consequences of their actions. It is not enough to simply be able to pick things up or push them; the robot must understand the value of the information it gains versus the cost of the change it creates. By teaching robots to weigh these factors, the researchers have taken a step toward machines that can navigate our cluttered world with the same care and adaptability that humans use. The system does not just see the world; it understands the difference between a helpful interaction and a disruptive one, allowing it to learn about its environment without destroying the very order it is trying to map.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.