← Latest papers
💻 computer science

Data Scaling Laws in Imitation Learning for Robotic Manipulation

This paper demonstrates that in robotic manipulation, generalization performance follows power-law scaling with the diversity of training environments and objects rather than the sheer volume of demonstrations, enabling highly effective zero-shot policies through a focused data collection strategy.

Original authors: Fanqi Lin, Yingdong Hu, Pingyue Sheng, Chuan Wen, Jiacheng You, Yang Gao

Published 2026-06-29
📖 4 min read☕ Coffee break read

Original authors: Fanqi Lin, Yingdong Hu, Pingyue Sheng, Chuan Wen, Jiacheng You, Yang Gao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to do a specific chore, like pouring water from a bottle into a cup. In the past, you might have thought the best way to teach it was to show the robot the exact same action thousands of times in the exact same kitchen. You'd think, "If I just give it enough practice, it will get perfect."

This paper says: No, that's not the most efficient way.

Instead of practicing the same move over and over, you should teach the robot to do that same move in many different kitchens with many different bottles.

Here is the breakdown of their findings using simple analogies:

1. The "Variety" vs. "Volume" Rule

The researchers discovered a surprising law about how robots learn. They found that variety is far more important than volume.

  • The Old Way (Volume): Showing a robot 1,000 videos of pouring water in one specific kitchen with one specific blue bottle.
  • The New Way (Variety): Showing a robot just 50 videos of pouring water, but in 50 different kitchens, using 50 different types of bottles (glass, plastic, tall, short, red, clear).

The Analogy: Think of it like learning to drive.

  • If you only practice driving in your driveway on a sunny day (high volume, low variety), you will crash the moment you hit a rainy street or a busy intersection.
  • If you practice driving in 32 different neighborhoods, in rain, snow, and traffic, but only drive for a short time in each (high variety, lower volume), you will become a much better driver much faster.

The paper found that once you have enough different scenarios (about 32 different environments), adding more practice videos for the same scenario barely helps at all.

2. The "Power Law" of Learning

The researchers found that the robot's ability to handle new, unseen situations follows a predictable mathematical curve called a Power Law.

The Analogy: Imagine the robot's "generalization skill" is a muscle.

  • If you lift a weight (add a new environment or object) once, your muscle grows a little.
  • If you lift weights in 2 different gyms, it grows more.
  • If you lift weights in 32 different gyms, your muscle becomes incredibly strong.
  • The paper shows that this growth isn't random; it follows a smooth, predictable line. The more different places and objects you train on, the better the robot gets at handling new places and objects it has never seen before.

3. The "One Afternoon" Breakthrough

The most exciting part of the paper is how efficient this method is.

The researchers tested this on four different tasks: pouring water, arranging a computer mouse, folding towels, and unplugging a charger.

  • They used four people (data collectors).
  • They worked for just one afternoon.
  • They collected data in 32 different environments with 50 demonstrations for each.

The Result: They trained robots that could perform these tasks with 90% success rates in completely new rooms with objects they had never seen before. They didn't need to fine-tune the robot or retrain it; it just worked immediately (zero-shot).

4. What About the Robot's "Brain"?

The paper also looked at the size of the robot's "brain" (the AI model).

  • The Visual Encoder (The Eyes): Making the robot's "eyes" bigger and smarter (using a larger model) definitely helped it see better and perform better.
  • The Action Model (The Hands): Surprisingly, making the "hand" part of the brain bigger didn't help. A smaller, simpler model was just as good as a giant one for these specific tasks. It seems the "hands" didn't need to be more complex; they just needed to have seen more variety.

The Bottom Line

The paper argues that to build a robot that can handle the real world, we shouldn't just dump massive amounts of repetitive data into it. Instead, we should focus on diversity.

If you want a robot that can pour water into any cup in any house, don't show it 10,000 videos of pouring in one kitchen. Show it 1,600 videos (32 locations × 50 tries) across 32 different kitchens with 32 different cups. That small, diverse dataset is the "secret sauce" to making robots truly adaptable.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →