← Latest papers
💻 computer science

Cloak: Zero-Shot Cross-Embodiment Manipulation by Masking the End-Effector from the VLA

The paper introduces Cloak, a training method that enables Vision-Language-Action models to achieve zero-shot cross-embodiment manipulation by masking the end-effector from wrist camera views during training, thereby allowing a model trained on a single robot to generalize to unseen hardware without collecting new data.

Original authors: Michael Piseno, Guy Tevet, C. Karen Liu

Published 2026-06-23
📖 4 min read☕ Coffee break read

Original authors: Michael Piseno, Guy Tevet, C. Karen Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot how to pick up a cup. You show it a video from the robot's own "eye" (a camera on its wrist). The problem is, in that video, the robot's own hand (the gripper) takes up a huge chunk of the screen.

If you train a robot using a specific type of hand—say, a simple two-fingered pincer—it learns to recognize that specific shape. If you then swap that robot's hand for a completely different one, like a human-like hand with five fingers, the robot gets confused. It's like if you taught someone to drive a car by only showing them the steering wheel of a Ford. If you suddenly put them in a Tesla with a different steering wheel, they might panic because the "visual map" in their brain doesn't match what they see.

The Problem: The "Hand" Distraction
In the world of robotics, collecting data is expensive and slow. We have huge libraries of data showing robots with two-fingered grippers doing tasks. But the future of robotics is moving toward complex, multi-fingered hands. We want to use our old data to teach these new hands, but the old robot's hand looks nothing like the new one. The robot's brain is too focused on the specific shape of the old hand to understand the new one.

The Solution: "Cloak"
The paper introduces a clever trick called Cloak. Think of it as a digital "blindfold" or a "muzzle" for the robot's vision.

  1. Hiding the Hand: Instead of letting the robot see its own hand in the camera feed, Cloak digitally paints over the hand with a mask. It's like putting a piece of tape over the part of the camera lens where the hand appears.
  2. Learning the Scene, Not the Hand: By hiding the hand, the robot is forced to look at the rest of the world—the table, the cup, the obstacles. It learns the logic of the task (e.g., "move forward until I hit the cup") without getting distracted by the specific shape of its own fingers.
  3. The "Chameleon" Training: To make sure the robot doesn't just learn to ignore a specific shape of tape, the researchers randomly change the shape of the "blindfold" during training. Sometimes the mask is big, sometimes small, sometimes weirdly shaped. This teaches the robot that "my hand might look different tomorrow, so I shouldn't rely on seeing it."

How It Works in Real Life
When the robot is actually doing the job:

  • The Eyes: The camera feed still has the hand, but the software instantly covers the hand with a digital mask before the robot's brain sees it.
  • The Brain: The robot's brain (a Vision-Language-Action model) sees a masked image and a text command like "pick up the cup." It figures out the movement based on the environment.
  • The Body: Once the brain decides "move my fingertips to this spot," a translator (called tip-pose retargeting) converts that instruction into the specific muscle movements needed for the new hand. If the new hand has five fingers, the system figures out how to move those five fingers to match the two-fingered plan.

The Results
The researchers tested this by training a robot on a simple two-fingered gripper using existing data. Then, they tried to use that same trained brain on three completely different robots it had never seen before:

  1. A different two-fingered gripper (different color and shape).
  2. A different robot arm entirely.
  3. A complex, five-fingered human-like hand.

The Outcome:

  • Without Cloak: The robot failed miserably on the new hands. It was confused because the visual "map" didn't match.
  • With Cloak: The robot performed almost as well on the new hands as it did on the original one. It successfully transferred its skills zero-shot (meaning it didn't need any new training data for the new hands).

The Big Takeaway
Cloak proves that you can separate a robot's "brain" (the skill of doing a task) from its "body" (the specific hardware). By hiding the body from the brain's view during learning, the data collected on one robot can be reused on completely different robots in the future. It's like teaching a person to swim by having them wear a blindfold so they focus on the water's movement rather than their own limbs; once they learn the rhythm of the water, they can swim with any kind of fins or flippers.

Limitations
The paper notes this works best for tasks that can be done by placing two points (like two fingertips) on an object. It doesn't yet solve complex tasks that require intricate in-hand manipulation (like juggling a ball inside a hand), as those rely heavily on the specific feel and shape of the fingers.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →