← Latest papers
💻 computer science

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

Open-AoE is an open, community-driven dataset and toolchain that bridges the gap between scalable smartphone-based egocentric video capture and embodied AI training by providing 2,000 hours of structured manipulation data alongside a comprehensive pipeline for processing, retargeting, and training various robot learning models.

Original authors: Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, Zitong Shan, Zhenchao Jin, Jiadong Hong, Taowen Wang, Yushi Feng, You Liu, Yibo W
Published 2026-07-21
📖 3 min read☕ Coffee break read

Original authors: Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, Zitong Shan, Zhenchao Jin, Jiadong Hong, Taowen Wang, Yushi Feng, You Liu, Yibo Wang, Yifan Yang, Zhaowen Zhou, Man Luo, Hao Cheng, Bo Zhang, Jianshu Li, Jiansheng Cai, Guocai Yao, Jize Zhang, Chenhao Lin, Renjing Xu, Lequan Yu, Chao Shen, Chunhua Shen, Zhe Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to teach a robot how to make a sandwich, fix a leaky faucet, or fold a shirt. You can't just feed it a library of textbooks; it needs to see and feel how humans do these things. This is the world of embodied intelligence, where computers learn by interacting with the physical world, just like we do. But there's a catch: robots see the world differently than we do. If you record a video from a camera on a shelf (a "third-person" view), the robot sees the whole room but misses the tiny finger movements. If you record from a robot's own eyes ("egocentric" view), it gets the right perspective, but getting enough high-quality data to teach it is incredibly hard and expensive. Usually, you need special, pricey cameras strapped to people's heads, or you have to manually label every single second of video to tell the computer what is happening. Without a massive, easy-to-use library of these "human-eye" videos, robots struggle to learn the subtle tricks of daily life.

This is where a new project called Open-AoE steps in, acting like a massive, open-source toolkit for teaching robots. The researchers behind it realized that everyone already has the perfect camera in their pocket: a smartphone. So, they built a system that turns thousands of ordinary people into data collectors. They created an app that records videos of people doing everyday tasks—like pouring coffee or typing on a keyboard—using their own phones. But they didn't just stop at recording; they built a giant, automated factory in the cloud that takes those raw, messy videos and turns them into super-organized, robot-ready lessons.

The team gathered a staggering 2,000 hours of these videos, collected by over 500 different people using more than 400 different types of smartphones. That's a lot of time! The videos cover over 400 different scenes (from kitchens to offices) and 8,000+ different tasks. What makes this special isn't just the volume, but the "magic" the system adds to the footage. Using advanced AI, the system automatically figures out exactly where the hands are moving (down to the 21 joints of the fingers), where the camera is moving, and breaks the video into tiny, meaningful chunks of action. It even writes text descriptions for what's happening.

Think of it like this: before, if you wanted to teach a robot, you had to hire a film crew, buy expensive gear, and spend years editing the footage. Open-AoE is like handing everyone a smartphone, saying "Go film your day," and then having a magical robot editor instantly turn those films into a perfect textbook for a robot to learn from. The paper shows that this approach works, providing a rich, diverse dataset that helps robots understand not just what to do, but how to move their hands and eyes to do it.

The researchers also built a set of tools (a "toolchain") that lets scientists take this data and do cool things with it. They can visualize the 3D movement of hands, "retarget" human movements to fit different robot bodies (like turning a human arm into a robot arm), and train different types of AI models. They found that by using this diverse, smartphone-based data, they can create a much stronger foundation for robots to learn from the real world, without needing a million dollars in equipment. It's a big step toward making robots that can actually help us with our daily chores, because now they have a massive, open library of human experiences to study.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →