← Latest papers
💻 computer science

ContactMimic: Humanoid Object Interaction via Contact Control

ContactMimic is a learning framework that enhances humanoid object interaction by explicitly tracking binary contact commands alongside keypoint trajectories, enabling precise physical contact and on-demand contact controllability across diverse manipulation tasks.

Original authors: Xinyao Li, Xialin He, Runpei Dong, Saurabh Gupta

Published 2026-07-10
📖 6 min read🧠 Deep dive

Original authors: Xinyao Li, Xialin He, Runpei Dong, Saurabh Gupta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to dance. For a long time, the best way to do this was to give the robot a video of a human dancer and say, "Copy these exact body positions." The robot would move its joints to match the human's hips, elbows, and knees perfectly.

But here's the problem: Just copying the shape of the dance isn't enough.

If a human dancer waves their hand near a whiteboard, they are just waving. If they press their hand against the whiteboard, they are wiping it. If a robot only looks at the hand's position, it might wave right next to the board and think it's doing a great job, even though it's not actually cleaning anything. It's like trying to open a door by standing right next to it without ever touching the handle.

This is the big idea behind a new project called CONTACTMIMIC. The researchers found that to make a robot truly useful for tasks like sitting, wiping, or pushing furniture, you can't just tell it where to be; you have to tell it whether to touch things.

The "Magic Switch" for Touching

The team built a system where, instead of just giving the robot a path to follow, they give it a switch. This switch is a simple "Yes/No" label for every part of the robot's body.

  • Switch ON (Contact ✔): "Okay, robot, I want you to press your hand against that whiteboard."
  • Switch OFF (Contact ✘): "Okay, robot, move your hand to the exact same spot, but don't touch the board. Just hover."

In their tests, they showed that with this switch, the same robot could do the exact same dance moves but change its behavior completely. It could sit in a chair and lean back, or sit in the same chair and stay upright without touching the backrest. It could pick up a box, or just hold its hands in the shape of holding a box without actually lifting it.

Why Was This So Hard?

You might think, "Why not just teach the robot to touch things naturally?" The researchers discovered a sneaky trap in the data.

When humans interact with objects, their body positions and their touches are usually locked together. If a human sits in a chair, their bottom is always touching the seat. If a human wipes a board, their hand is always touching the surface. Because of this, if you just show a robot thousands of videos of people sitting, the robot learns: "Oh, if the hips are in this position, the bottom must be touching the chair." It stops listening to the "touch" command because it thinks the position tells the whole story.

To fix this, the team had to get creative. They took their training data and started breaking the rules:

  1. The "Ghost" Object: They took a video of someone wiping a board but told the robot, "The board isn't there." The robot had to learn to move its hand to the board's spot without touching anything.
  2. The "Inflated" Object: They made the objects in the simulation bigger (like inflating a balloon) so the robot had to move its hand around the object to avoid hitting it, even though the "touch" command said it should be touching.
  3. The "Fake" Touch: They took a video where the robot was supposed to touch, but they told it, "Don't touch," and vice versa.

By mixing these up, they forced the robot to stop guessing based on position and start actually listening to the "touch" switch.

Did It Work?

The team tested this on a real robot called the Unitree G1 (a 29-joint humanoid) and in a computer simulation.

  • In the Simulation: They ran 10 different tasks, like kicking a chair, stepping on a chair, and picking up a box. When they flipped the switch to "Contact," the robot made the contact. When they flipped it to "No Contact," it stopped touching. The robot could even lift a box just by being told to touch it, without needing any special "lift the box" instructions.
  • In the Real World: They tested five real-life motions, including wiping a whiteboard and leaning on a chair back.
    • When told to wipe, the robot's hand pressed hard enough to leave a visible mark on the board.
    • When told to hover, the hand moved to the same spot but didn't leave a mark.
    • When told to lean, the robot rested its weight on the chair back. When told to not lean, it stayed upright, balancing on its own.

The results were impressive. In the real world, the robot succeeded in following the contact command about 90% to 100% of the time for most tasks (for example, 9 out of 10 times for leaning on a backrest, and 5 out of 5 for wiping the board).

What It Can't Do (Yet)

The paper is very clear about what this isn't.

  • It's not a "universal" robot that can do any task with one brain. The researchers trained a separate "brain" (policy) for each specific motion. They haven't figured out how to make one single robot that can do all these things at once yet.
  • It relies on high-quality videos of humans interacting with objects. They didn't use random videos from the internet; they used a specific dataset called HUMOTO.
  • They didn't use special touch sensors on the robot's skin. Instead, the robot figured out if it was touching something just by feeling its own muscles and joints (proprioception). They proved this works by showing that if you asked a simple computer program to guess "Is the robot touching?" based on the robot's internal feelings, it was right almost all the time (with F1 scores around 0.90 to 0.99 for most tasks).

The Bottom Line

The paper suggests that if you want a robot to do useful, physical jobs, you have to stop just telling it where to go. You have to explicitly tell it whether to touch. By breaking the link between "position" and "touch" during training, they taught the robot to listen to a simple switch. It's a step toward robots that don't just look like they are doing a task, but actually do the task by making the right physical contact.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →