Compliance for Free: Learning Identifiable Impedance via Bilateral Teleoperation
This paper proposes a method using four-channel bilateral teleoperation to identify and record operator stiffness from existing joint-torque sensors, enabling the fine-tuning of vision-language-action models to generate direction-dependent compliance labels at zero annotation cost for contact-rich tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Robots have become remarkably good at seeing the world and deciding where to move next. Modern artificial intelligence can look at a picture of a messy room and tell a robot arm exactly how to pick up a cup or open a drawer. But there is a final, critical step where these machines often stumble: the moment they touch something. While a robot can be told where to go, it struggles to know how hard to push. Tasks that require a gentle touch, like wiping a whiteboard without scratching it, or inserting a peg into a tight hole, depend on a property called compliance. This is the ability to yield or resist pressure in a controlled way. For a long time, teaching robots this skill has been difficult because the standard ways humans show robots what to do only record the robot's final position. They miss the invisible force the human was applying, leaving the robot to guess whether it should be stiff or soft.
A team of researchers in Barcelona has found a way to solve this missing piece of information without adding expensive force-torque or tactile sensors. They discovered that by using a specific type of remote control setup, they could mathematically separate the human's intended target from the force they applied using only the joint-torque sensors already present on the robot arm. In their experiments, a human operator controlled a robot arm through a system where the operator's own arm moved a second, identical robot arm. By watching how the operator's arm moved compared to how the robot arm actually moved, the researchers could calculate exactly how stiff or soft the operator wanted the robot to be at every single moment. They used this new data to train an artificial intelligence model that can now take a simple spoken instruction, such as "wipe this mark firmly," and adjust its own stiffness to match, rather than just moving to a location.
The core of the problem the researchers tackled is a confusion that happens whenever we try to teach a robot by watching a human. Imagine a human pushing a robot arm against a wall. If the robot ends up in a certain spot, we do not know if the human pushed hard against a stiff spring or gently against a soft spring; both scenarios could result in the robot ending up in the exact same place. Standard recording devices only see the final position, so they cannot tell the difference. This makes it impossible to teach a robot the concept of "firm" versus "gentle" just by watching where the robot goes. The researchers realized that to fix this, they needed a second, independent measurement of where the human intended to go, distinct from where the robot actually ended up.
To solve this, they used a four-channel bilateral teleoperation system. In this setup, a human holds a "leader" robot arm, and a "follower" robot arm tries to copy it. Crucially, the system does not just send position commands; it also sends force information back and forth. When the follower arm hits an obstacle, the human feels that resistance in their own hand. Because the human is actively feeling the contact, their own arm's movement represents their true intention, while the follower's movement shows how the robot reacted to the environment. By comparing the leader's path with the follower's path, the researchers could calculate the difference between where the human wanted to be and where the robot actually was. This difference, combined with the force the robot felt (which was estimated using the robot's native joint-torque sensors), allowed them to work backward and determine the exact stiffness the human was applying.
Using this method, the team collected a large set of demonstrations where humans wiped a whiteboard. They instructed the humans to wipe marks on the board either "normally" or "firmly." Because of their new calculation method, they could extract a precise label for the stiffness used in every single moment of the wiping motion, without needing any special force sensors on the robot's wrist. They then used these labels to fine-tune a large vision-language model, a type of artificial intelligence that understands images and text. They taught this model to output not just a target position, but also a target stiffness value for every move it makes.
The results showed that this approach worked exactly as intended. When the researchers tested the new robot policy, they found that it could genuinely change how hard it pressed based on the language instruction. When told to wipe "firmly," the robot increased its contact force to an average of 9.1 Newtons. When told to wipe "normally," it reduced that force to 6.4 Newtons. This change was statistically significant and consistent. In contrast, other robot policies tested in the same study failed to change their behavior based on the instruction. Some policies that were given force data as an input still moved too rigidly, while others that tried to guess the stiffness based on fixed rules could not adapt at all. Only the policy trained with the new, extracted labels successfully linked the word "firmly" to a physical increase in pressure.
The researchers also verified that the robot learned to be selectively stiff. It did not just become uniformly harder in all directions; it became stiff in the direction of the wiping motion to resist being pushed off course, while remaining softer in other directions. This anisotropic behavior, or direction-dependent stiffness, is a sophisticated trait that mimics how a human hand naturally adjusts to a task. The robot achieved a 50 percent success rate in removing ink from the board, which was the highest among the five different policies tested, and it did so while triggering fewer safety stops caused by excessive force.
Despite these successes, the study acknowledges certain limitations. The method relies on the robot having sensors that measure the torque in its joints, which are common in research robots but not always in cheaper commercial models. The technique also required the robot to stay in a specific posture during calibration to ensure the force calculations remained accurate. Furthermore, the rotational stiffness, or how the robot resists twisting, was harder to measure accurately because the wiping task did not involve enough twisting motion to generate clear data. The researchers noted that while their method works well for this specific task, it is not a universal solution for every type of contact a robot might encounter.
The significance of this work lies in how it bypasses the need for expensive, specialized force-torque or tactile hardware to teach robots about touch. By using the existing joint-torque sensors on a robot arm and a clever mathematical approach to interpret human intent, the researchers created a way to generate high-quality training data for compliance at zero extra cost. This suggests that robots could learn to handle delicate or variable-force tasks much more effectively in the future, simply by observing humans through a system that captures both movement and the feeling of resistance. The study demonstrates that the key to better robot touch may not be better sensors, but better ways of reading the signals we already have.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.