← Latest papers
💻 computer science

EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning

EgoInfinity is a modular, web-scale 4D data engine that transforms arbitrary internet videos into metric hand-object interaction representations and executable robot trajectories, enabling scalable, human-in-the-loop-free robot learning and retargeting across diverse morphologies and viewpoints.

Original authors: Gaotian Wang, Kejia Ren, Andrew Morgan, Yiting Chen, Howard H. Qian, Podshara Chanrungmaneekul, Kaiyu Hang

Published 2026-06-17
📖 5 min read🧠 Deep dive

Original authors: Gaotian Wang, Kejia Ren, Andrew Morgan, Yiting Chen, Howard H. Qian, Podshara Chanrungmaneekul, Kaiyu Hang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a massive library of YouTube videos showing people doing everyday tasks—opening boxes, pouring soup, cutting fruit, or wiping tables. Right now, robots can't really "learn" from these videos because the footage is just flat pictures (2D). The robot doesn't know how heavy the box is, how far away the soup is, or exactly how the human's fingers touched the handle. It's like trying to learn how to drive a car just by watching a black-and-white movie; you can see the movement, but you can't feel the steering wheel or the road.

EgoInfinity is a new "magic engine" that turns those flat, messy internet videos into a 3D, physics-ready instruction manual for robots. Here is how it works, broken down into simple steps:

1. The "Detective" Phase (Finding the Clues)

First, the engine scans through millions of video clips (specifically from a huge dataset called Action100M). It acts like a detective looking for the good stuff. It ignores boring parts and focuses only on the moments where hands are actually doing something.

  • The Metaphor: Think of it like a film editor who quickly scans 100 hours of raw footage to find the 5 minutes where the actor actually speaks, ignoring all the silence.

2. The "3D Translator" (Turning Pictures into Geometry)

Once it finds a good clip, the engine starts translating the 2D video into 3D data. It uses smart AI tools to guess:

  • The Shape: What does the object look like in 3D? (Is it a round apple or a flat box?)
  • The Position: Where is the object in space?
  • The Hands: How are the human fingers moving?
  • The Scale: How big is everything? (Is that cup 2 inches tall or 2 feet tall?)

The paper calls this "metric calibration."

  • The Metaphor: Imagine looking at a photo of a person holding a coffee cup. A normal person might guess the size. EgoInfinity is like a super-accurate 3D scanner that instantly knows the cup is exactly 4 inches tall and the hand is 12 inches away, turning the flat photo into a virtual 3D model you could walk around.

3. The "Reality Check" (Fixing the Mistakes)

AI sometimes gets confused. If a hand covers an object, the AI might lose track of where the object is. EgoInfinity has a special "Interaction-Aware Refinement" step. It looks at the relationship between the hand and the object.

  • The Logic: If the hand is holding the object, the object must move with the hand. If the hand is still, the object shouldn't be floating away.
  • The Metaphor: It's like a dance instructor watching a video. If the video glitches and the dancer's partner suddenly floats into the air, the instructor says, "No, that's wrong. If they are holding hands, they must move together." The engine fixes these glitches to make the physics look real.

4. The "Universal Adapter" (Teaching Different Robots)

This is the most clever part. The engine doesn't just copy the human's body movements exactly. Humans have long arms and two hands; robots might have short arms, wheels, or grippers that look nothing like hands.

  • The Solution: EgoInfinity figures out the goal of the movement (e.g., "pick up the cup") and then translates that goal into the robot's own language. It asks, "How can this specific robot achieve the same result?"
  • The Metaphor: Imagine you are teaching a dog to fetch a ball. You don't tell the dog to "run like a human." You tell the dog, "Go get the ball," and the dog figures out how to use its paws and legs to do it. EgoInfinity does this for robots: it takes the human's "fetch" and tells a robot arm how to "fetch" using its own joints.

What Did They Actually Build?

The authors didn't just write a theory; they built a working system and tested it:

  • A Web Engine: They created a tool that can process these videos automatically without humans needing to label every frame.
  • A Dataset: They processed 106 videos from the Action100M collection, creating a library of 3D hand-and-object data.
  • Real Robot Tests: They successfully taught real robots (like a Unitree G1 robot and a dual-arm Franka robot) to perform tasks like cutting, pouring, and wiping just by watching these processed videos. They also trained a robot hand to grasp different objects (apples, bananas, soup cans) using the data as a guide.

What Are the Limits?

The paper is honest about what it can't do yet:

  • Camera Movement: It works best when the camera is mostly still (like a tripod). It struggles if the camera is shaking wildly or held in a moving hand.
  • Touch: It can't feel things. It knows where the hand is, but it doesn't know how hard the robot is squeezing or if the object is slippery.
  • Perfect Precision: It's great for general tasks, but it might not be precise enough for delicate surgery or tasks requiring exact fingertip placement.

In short: EgoInfinity is a bridge. It takes the messy, unstructured world of human YouTube videos and turns them into clean, 3D, physics-accurate instructions that robots can actually read and follow, allowing them to learn new skills without needing expensive sensors or human teachers to record every move.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →