← Latest papers
🤖 AI

ATOM-Bench: A Real-World Benchmark for Atomic Skills and Compositional Generalization in Manipulation Policies

The paper introduces ATOM-Bench, a real-world benchmark featuring 30 atomic and 24 compositional manipulation tasks across single- and dual-arm robots, which evaluates generalist policies by distinguishing between failures in fine-grained motor execution and limited compositional generalization through extensive physical rollouts.

Original authors: Zenan Wu, Bingqing Wei, Lu Liu, Zheqi He, Xi Wang, Jiakang Liu, Zehui Li, Guocai Yao, Jing-Shu Zheng, Xi Yang, Yongtao Wang

Published 2026-06-16
📖 4 min read☕ Coffee break read

Original authors: Zenan Wu, Bingqing Wei, Lu Liu, Zheqi He, Xi Wang, Jiakang Liu, Zehui Li, Guocai Yao, Jing-Shu Zheng, Xi Yang, Yongtao Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to be a helpful kitchen assistant. You show it how to pick up a cup, pour water, and put it in the sink. The robot seems to learn quickly. But then, you ask it to "pick up the red cup and pour it into the small bowl." Suddenly, the robot freezes or spills everything.

This is the problem ATOM-Bench was built to solve.

The Problem: The "Cookbook" vs. The "Chef"

Current robot "brains" (called manipulation policies) are often like students who have memorized a specific cookbook. They can follow the exact steps you showed them perfectly. But they haven't learned the ingredients or the techniques on their own. If you ask them to cook a dish you didn't show them in the book, they fail because they can't mix and match what they know.

The researchers wanted to know: Does the robot actually understand the basic moves (like "pick up" or "pour"), or is it just memorizing the whole recipe?

The Solution: ATOM-Bench (The "Atomic" Test)

The team created a new testing ground called ATOM-Bench. Think of this as a gym for robots, but instead of just running laps, they break every task down into its smallest, indivisible parts, which they call "atoms."

They split these atoms into two types:

  1. Motor Atoms (The Hands): These are the physical moves. Pick up, pour, push, stack, open a drawer.
  2. Instruction Atoms (The Brain): These are the logic rules in the command. The "red" one, the "biggest" one, "exactly two" of them, or "not the blue one."

How the Test Works

The researchers set up a real-world lab with two types of robots:

  • One-armed robots (like a single human arm).
  • Two-armed robots (like a human with two hands working together).

They taught the robots 30 basic "atomic" tasks using 3,000 human demonstrations (like showing a child how to do it 100 times). Then, they gave the robots 24 brand-new, "hidden" tasks that mixed these atoms together in ways the robots had never seen before.

For example, if the robot learned "pour" (Motor) and "red" (Instruction) separately, the test would ask: "Pour the red beans."

The Results: Good Hands, Confused Brains

After testing five different advanced robot models, the researchers found some surprising things:

  • The "Simple" Stuff Works: The robots are actually pretty good at following simple instructions like "pick up the blue block." They can handle basic visual cues.
  • The "Hard" Stuff Fails: When the task gets tricky—like pouring liquid without spilling, counting exactly two items, or figuring out which object is "not red"—the robots struggle. Their hands aren't steady enough, and their brains can't do the math.
  • The "Mix-and-Match" Problem: This is the biggest finding. Even when a robot is great at the individual parts (like pouring and recognizing red), it fails miserably when asked to combine them into a new task. It's like a musician who can play a perfect scale and a perfect chord, but can't play a song that uses both.

The New Scorecards

To understand why the robots failed, the authors invented two new scorecards:

  1. Atomic Score (AS): Did the robot know how to do the individual pieces? (e.g., Did it know how to pour?)
  2. Compositional Failure Share (CFS): Did the robot fail because it didn't know the pieces, or because it couldn't put them together?

The results showed that for the best robots, the failure wasn't that they didn't know how to pour; it was that they couldn't combine the pouring skill with the "red" instruction to solve the new puzzle.

The Takeaway

The paper concludes that while we are building impressive robot models, they are still far from being truly "general" helpers. They are currently more like parrots that repeat what they've heard than chefs who understand how to cook. They can do the specific moves they were trained on, but they haven't yet learned how to creatively recombine those moves to handle the messy, unpredictable real world.

ATOM-Bench is now a tool for scientists to diagnose exactly where the robot is breaking down: Is it the hands? The eyes? Or the ability to think ahead?

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →