← Latest papers
💻 computer science

FOCI Policy: Focus on Object-Centric Interactions for Relational Manipulation Policies

The paper proposes the FOCI Policy, an object-centric framework that improves generalization and data efficiency in rigid relational manipulation by abstracting skills into temporally compact interaction segments and spatially invariant relative SE(3)SE(3) motions between task-relevant objects.

Original authors: Ze Fu, Pinhao Song, Yutong Hu, Renaud Detry

Published 2026-09-09
📖 5 min read🧠 Deep dive

Original authors: Ze Fu, Pinhao Song, Yutong Hu, Renaud Detry

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Robots have long struggled with the simple act of moving objects from one place to another. While they can be programmed to perform repetitive tasks in a factory, teaching them to handle new situations with a single glance or a brief demonstration remains a profound challenge. Most current approaches try to teach a robot by showing it exactly how to move its own arms and fingers, essentially drilling a form of muscle memory. This method requires vast amounts of data, often hundreds of demonstrations, to cover every possible way an object might be positioned or a room might be arranged. If the robot encounters a slight change, such as a cup being moved a few inches to the left, the rigid instructions often fail. Researchers are now exploring a different path: instead of focusing on the robot's body, they are teaching machines to focus on the relationship between the objects themselves. By understanding how a toothbrush moves relative to a cup, rather than how the robot's gripper moves through space, a system might learn to generalize skills with far less effort.

A team of researchers at KU Leuven has developed a new framework called FOCI Policy that puts this idea into practice. Their work suggests that many complex tasks are governed by short, critical moments where objects interact, while the rest of the movement is merely travel. Imagine trying to learn how to pour a glass of water. The act of lifting the pitcher and walking to the table can be done in countless ways, but the moment the water flows from the spout into the glass is tightly constrained and follows a specific pattern. The researchers observed that if a robot can isolate these brief, high-stakes interaction phases and ignore the noisy, variable travel time, it can learn the task much more efficiently. They built a system that automatically scans a demonstration video, identifies exactly when the objects begin to touch or influence one another, and extracts that specific segment.

Once this critical segment is isolated, the system does not try to memorize the robot's exact path. Instead, it learns the relative motion between the two objects involved. If a human demonstrates placing a wine bottle into a rack, the robot learns the movement of the bottle relative to the rack, not the movement of the robot's arm. This approach makes the skill independent of the robot's specific body shape or the exact layout of the room. The system can then take this learned relationship and apply it to a new situation, even if the objects are in different positions or the robot is a different model entirely. In their tests, the researchers trained the system on just a single demonstration for each task. When faced with new scenarios, the robot successfully performed complex actions like stacking wine bottles, screwing in nails, or pouring liquids, often outperforming other methods that were trained on significantly more data.

The results were particularly striking when the visual conditions changed. In one set of experiments, the researchers introduced severe visual disturbances, such as changing the lighting, adding distracting objects, or altering the texture of the items. While other advanced systems trained on a hundred demonstrations saw their success rates drop dramatically under these conditions, the new method maintained a level of performance that was comparable to those heavily trained models, despite using only one-tenth of the data. This suggests that by focusing on the geometric relationship between objects, the robot becomes less confused by visual noise. The system proved robust enough to handle real-world tasks on a physical robot arm, successfully completing tasks like sweeping dust or opening a drawer after seeing the action just once.

However, the researchers are careful to note that this approach is not a universal solution for every type of movement. The method works best for tasks involving rigid objects that have clear, distinct phases of interaction, such as inserting a peg into a hole or placing a book on a shelf. It is less effective for tasks that involve continuous contact, like pushing a soft object across a table, or for situations where the objects are so small or hidden that the robot cannot reliably determine their position. The system relies on the ability to see the objects clearly; if the robot cannot estimate where an object is, it cannot calculate the correct relative motion. Despite these limitations, the findings offer a compelling shift in how robots learn. By stripping away the unnecessary details of how a robot moves and focusing on the essential dance between objects, the researchers have shown that machines can acquire new skills with a fraction of the data previously thought necessary. This efficiency brings us closer to a future where robots can learn new tasks in minutes rather than days, adapting quickly to the unpredictable nature of the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →