Vision-Force Admittance Learning for Peg Insertion into a Movable Hole
This paper proposes a Vision-Force Admittance Learning (VFAL) framework that fuses asynchronous visual feedback from foundation models with high-frequency force control to enable robust, millimeter-precision peg-in-hole insertion into moving holes in dynamic environments.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the quiet, controlled world of a factory floor, a robot arm can be a marvel of precision. It can pick up a screw and drive it into a hole with a steadiness no human hand can match, provided the hole does not move. This is the realm of static manipulation, where the target sits still, and the robot's challenge is simply to find its way to a fixed point. But the real world is rarely so still. Imagine a car part moving along a conveyor belt, or a robot mounted on a vehicle that is driving over uneven ground. In these dynamic environments, the target is in motion, and the robot must not only find the hole but chase it, all while maintaining the millimeter-level accuracy required to fit a peg into a tight space. This is a fundamental conflict: the need for extreme precision clashes with the chaos of a moving target. Traditional robots struggle here because they rely on either slow, high-level vision to see where the hole is, or fast, local touch sensors to feel the moment of contact. Vision is too slow to react to sudden shifts, while touch alone cannot see far enough ahead to plan a smooth path.
Researchers at New York University and General Motors have developed a new way to solve this problem, creating a system that allows a robot to insert a peg into a moving hole with remarkable success. They call their method Vision-Force Admittance Learning. Instead of forcing the robot to choose between seeing and feeling, this system lets the two senses work together in a specific, asynchronous rhythm. The robot uses a camera to estimate where the hole is, but because cameras are slow, this information is treated as a general guide rather than a strict command. Simultaneously, the robot relies on a force sensor at its wrist, which reacts hundreds of times per second to the physical pressure of the peg touching the hole. The innovation lies in how these two streams of data are combined. The fast force sensor makes the immediate, split-second adjustments needed to keep the peg aligned, while the slower camera updates the robot's understanding of the hole's overall position. This visual update acts as a gentle correction, or a regularizer, that keeps the robot from drifting too far off course, even as the force sensor handles the rapid, chaotic vibrations of the insertion.
To test this, the team built a real-world experiment involving a robotic arm holding a peg and a base with a hole that could move. They used a special type of fastener, a pushnut with flexible fins, which is common in industrial assembly. The setup included a camera to watch the scene and a force sensor to feel the contact. They tested the robot in three different scenarios: a hole that stayed still, a hole moving in a straight line, and a hole moving in three dimensions on a motorized slider. The results were striking. When the hole was moving, a robot relying only on vision failed to insert the peg most of the time, while a robot relying only on force sensors also struggled, often losing its way or applying too much force. However, the new system, which fused the two senses, achieved a success rate of 80 percent in the most difficult three-dimensional moving scenario. In the straight-line moving scenario, it succeeded every single time.
The researchers found that the key to this success was not just having both sensors, but how they were timed. They compared their method to a version where the vision and force data were forced to wait for each other, a process that slowed everything down. By letting the fast force sensor drive the motion while the slower vision system updated the goal in the background, the robot could react instantly to physical contact without losing its global sense of direction. They also discovered that the system could recover from mistakes. If the camera lost track of the hole, the robot would safely back away and wait for the vision to return. If the peg got stuck or the force readings indicated a bad angle, the robot could lift the peg and try again. This ability to adapt and recover made the system robust enough to handle the unpredictable nature of a moving target.
The study also explored how different settings affected the outcome. They found that the balance between trusting the camera's position and the force sensor's feedback was critical. If the robot trusted the camera too much, it ignored the physical reality of the peg hitting the side of the hole. If it trusted the force sensor too much, it would drift aimlessly when the hole moved. By finding a middle ground, the robot learned to use the camera to stay on the right path and the force sensor to make the fine adjustments needed to slide the peg in. This approach worked even when the target moved in complex, non-linear ways, proving that the system could handle more than just simple, predictable motion.
Ultimately, this work demonstrates that robots can be taught to perform delicate tasks in environments that are not perfectly still. By combining the slow, broad view of a camera with the fast, sensitive touch of a force sensor, the researchers created a method that mimics how humans might handle a moving object: using our eyes to know where to go and our hands to feel the way. The system does not require the robot to be perfect from the start; it allows for errors and corrects them in real-time. While the current system assumes the robot has already picked up the peg, the ability to insert it into a moving hole is a significant step toward robots that can work alongside humans in dynamic, real-world settings, from assembly lines to mobile vehicles, without needing the environment to stop for them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.