Beyond Implicit Force: Evaluating Explicit Force-Torque Proxies in Action Chunking with Transformers
This paper demonstrates that while Action Chunking with Transformers (ACT) often relies on implicit teleoperation tracking errors to achieve contact-rich manipulation, replacing these hidden cues with explicit joint-torque proxies derived from onboard sensors effectively recovers or even enhances force-aware behavior across diverse real-world tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to do delicate chores, like wiping a table or plugging in a charger. You can't just show the robot a video; it needs to feel what it's doing. But robots don't have skin, and cameras can't see "pressure" or "friction." This is the tricky world of contact-rich manipulation, where a robot has to interact with the physical world without breaking things or getting stuck. To teach robots, scientists often use a method called Imitation Learning, where the robot watches a human do a task and tries to copy them. A popular tool for this is called ACT (Action Chunking with Transformers), which is like a super-smart predictor that guesses the next few moves a robot should make all at once, rather than one step at a time. The big question everyone is asking is: How does the robot know when to push harder or back off if it can't actually "feel" the object? Does it need expensive, specialized sensors, or can it figure it out just by watching?
This paper dives into a surprising secret hidden in how we teach robots. The researchers found that the popular ACT robot was actually "cheating" a little bit. When humans teach these robots using a special "leader-follower" setup (where a human moves a master arm, and the robot copies it with a slave arm), the robot wasn't just copying the position of the arm. It was secretly learning from the struggle. If the robot's arm hit a wall and couldn't move as fast as the human's arm, that tiny lag or "tracking error" acted like a hidden signal, telling the robot, "Hey, something is in the way, push harder!" The paper asks: What if we take away that hidden struggle signal? Does the robot forget how to handle contact?
To find out, the authors built a version of the robot that ignores the human's "struggle" and only predicts where the robot's own arm should go. They called this ACT-o. When they tested it on real-world tasks like wiping a board, plugging in a charger, or pressing a soft bottle, the robot completely fell apart. Without that hidden "lag" signal, the robot became hesitant, froze, or just hovered above the object, failing to apply the necessary force. It was like a musician who could play the notes but forgot how to press the keys hard enough to make a sound.
However, the story doesn't end in failure. The researchers then tried a simple fix: they gave the robot a new sense using data it already had. Instead of expensive external sensors, they used the robot's own internal motor current (how much electricity the motors are using) as a "torque proxy"—a rough estimate of how much force the robot is exerting. When they fed this simple "force feeling" back into the robot, it came back to life. The robot could suddenly distinguish between a hard surface and a soft one, stop exactly when it touched a block of foam, and even perform gentle "taps" to align a plug.
The paper suggests that while the hidden struggle from the old teaching method was a powerful, accidental teacher, we don't need to rely on it anymore. By simply adding a basic estimate of force (derived from motor current) to the robot's brain, we can make it just as good, or even better, at handling contact than before. This means we can teach robots to do delicate, force-sensitive jobs even on cheap hardware that doesn't have fancy, expensive force sensors, as long as we give them a way to "feel" their own motors working.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.