← Latest papers
🤖 machine learning

BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames

The paper proposes Big Picture Policies (BPP), a robot imitation learning approach that leverages vision-language models to select a minimal set of meaningful keyframes from history, thereby overcoming spurious correlations and distribution shift to achieve significantly higher success rates in long-context manipulation tasks compared to existing methods.

Original authors: Max Sobol Mark, Jacky Liang, Maria Attarian, Chuyuan Fu, Debidatta Dwibedi, Dhruv Shah, Aviral Kumar

Published 2026-02-19
📖 5 min read🧠 Deep dive

Original authors: Max Sobol Mark, Jacky Liang, Maria Attarian, Chuyuan Fu, Debidatta Dwibedi, Dhruv Shah, Aviral Kumar

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to make a cup of coffee.

If you just show the robot the current picture of the kitchen, it might get confused. "Is the coffee machine on? Did I already put the sugar in? Did I drop the spoon?" Without remembering what happened five seconds ago, the robot is like a goldfish with a 3-second memory. It keeps trying to add sugar even though the cup is full, or it keeps opening the same drawer over and over because it forgot it already checked it.

This is the problem with most robot brains today: they are memoryless.

But what if we just gave the robot a video camera that records everything it has ever seen? You might think, "Great! Now it has a perfect memory!"

Surprisingly, the paper says no, that makes it worse.

The Problem: The "Spurious Correlation" Trap

Imagine you are teaching a robot to find a hidden key in a house. You show it 100 videos of humans finding the key. In 99 of those videos, the human checks the kitchen before the living room.

If you give the robot a video of everything that happened, the robot might get lazy. Instead of learning "Check the kitchen, then check the living room," it learns a weird shortcut: "If I see a picture of the kitchen, I must be in the living room phase."

This is called a spurious correlation. The robot isn't understanding the task; it's just memorizing the specific order of events in the training videos.

When you send this robot into a real house where the human checked the living room first, the robot gets confused. It sees the living room, thinks "Oh, I must be done with the kitchen!" and stops working. It fails because the real world doesn't always follow the exact pattern of the training videos.

The space of all possible things a robot could see in a long task is exponentially huge. You can't possibly show the robot every single variation of "checking a drawer" or "dropping a spoon" during training. So, if you feed it raw history, it just memorizes the few examples it saw and fails when things look slightly different.

The Solution: BPP (Big Picture Policies)

The authors propose a clever fix called BPP.

Instead of feeding the robot a 30-minute video of "everything that happened," BPP asks a super-smart AI (a Vision-Language Model, or VLM) to act as a highlight reel editor.

Think of it like this:

  • Naïve Approach: You give the robot the entire raw footage of a movie. It gets overwhelmed and misses the plot.
  • BPP Approach: You give the robot a "Cliff's Notes" summary. The AI editor scans the video and only keeps the Keyframes.

What is a Keyframe?
A keyframe is a moment where something meaningful happened.

  • Not a keyframe: The robot's hand moving slowly through the air.
  • Keyframe: The robot successfully dropping a marshmallow into a bowl.
  • Keyframe: The robot opening a drawer and seeing it's empty.
  • Keyframe: The robot pressing the "Task Complete" button.

The robot only looks at these specific, important moments. It ignores the boring, repetitive stuff in between.

Why This Works (The Magic Analogy)

Imagine you are trying to learn how to drive a car by watching a video of a race.

  • The Naïve Method: You watch the video frame-by-frame. You memorize that "in this specific race, the car turned left at the red tree." But if you go to a different race where the tree is blue, you crash because you memorized the tree, not the turn.
  • The BPP Method: A coach watches the video and tells you: "Remember these three things: 1. Turn left at the intersection. 2. Brake at the curve. 3. Accelerate on the straight."

The coach (the VLM) filters out the noise. The robot learns the concept of the task (the "Big Picture") rather than the specific visual details of the training videos.

Because the robot only focuses on these few, important "milestones," it doesn't get confused by the millions of other ways the robot could have moved between those milestones. It generalizes much better.

The Results

The team tested this on real robots doing tricky tasks like:

  • Searching through drawers for a key.
  • Stacking puzzle pieces.
  • Putting marshmallows in a bowl.

The Results were amazing:

  • Robots with no memory failed miserably.
  • Robots with "raw memory" (watching everything) did slightly better but still failed often because they got confused by the noise.
  • The BPP robots were 70% more successful than the next best method. They could track their progress, realize when they made a mistake, and keep going without getting stuck in loops.

Summary

BPP is like giving a robot a smart summary of its own past instead of a raw video tape. By using an AI to pick out only the most important moments (the "Keyframes"), the robot learns the logic of the task rather than memorizing the accidents of the training data. This allows it to handle long, complicated jobs in the real world without getting lost in the details.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →