← Latest papers
💻 computer science

FrameSkip: Learning from Fewer but More Informative Frames in VLA Training

The paper introduces FrameSkip, a data-layer framework that enhances Vision-Language-Action (VLA) training efficiency and performance by selectively sampling high-importance frames based on action variation and visual coherence, achieving a 76.15% success rate across benchmarks while retaining only 20% of the original trajectory frames.

Original authors: Bin Yu, Shijie Lian, Xiaopeng Lin, Zhaolong Shen, Yuliang Wei, Changti Wu, Hang Yuan, Haishan Liu, Bailing Wang, Cong Huang, Kai Chen

Published 2026-05-14
📖 3 min read☕ Coffee break read

Original authors: Bin Yu, Shijie Lian, Xiaopeng Lin, Zhaolong Shen, Yuliang Wei, Changti Wu, Hang Yuan, Haishan Liu, Bailing Wang, Cong Huang, Kai Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to make a cup of coffee. You record a video of a human doing it from start to finish. This video is 10 minutes long.

The Old Way (The Problem)
In the past, when training AI robots, researchers would feed the computer every single frame of that 10-minute video. They treated every second as equally important.

But think about the video:

  • Minutes 0–2: The human walks slowly to the kitchen counter. (Nothing much happens).
  • Minute 3: They grab the coffee mug. (This is a critical moment!).
  • Minutes 4–8: They walk slowly to the coffee machine and wait. (Again, mostly just walking).
  • Minute 9: They press the button and the coffee pours. (Another critical moment!).
  • Minute 10: They walk away.

The paper argues that the "old way" is inefficient. The robot spends 80% of its study time watching the human walk slowly, and only 20% of its time watching the actual tricky parts (grabbing, pouring). It's like studying for a math test by reading the same boring chapter over and over, while only glancing at the actual formulas you need to solve the problems.

The New Solution: FRAMESKIP
The authors created a tool called FRAMESKIP. Think of it as a super-smart video editor that works before the robot starts learning.

Instead of showing the robot the whole 10-minute video, FRAMESKIP cuts out the boring parts and creates a "highlight reel." It keeps the slow walking parts very short but zooms in and repeats the important parts (like the hand grabbing the mug) so the robot sees them more often.

How does it know what to keep?
FRAMESKIP uses four simple "clues" to decide which frames are important:

  1. Action Changes: Did the robot's hand suddenly move fast? (Keep that frame).
  2. Visual Changes: Did the picture change because the object moved? (Keep that frame).
  3. Task Progress: Is this the middle of the task where things usually go wrong? (Keep that frame).
  4. Gripper Moves: Did the robot's "fingers" open or close? (Keep that frame).

The Magic Trick
The coolest part is that FRAMESKIP doesn't change the robot's brain or how it learns. It just changes what data the robot sees. It's like giving the same student a better study guide instead of trying to make the student smarter.

The Results
The researchers tested this on three different robot simulation games.

  • The Result: When they used the "highlight reel" (keeping only 20% of the original frames), the robots actually got better at their tasks than when they watched the full, boring video.
  • The Score: The robots succeeded about 76% of the time with the new method, compared to only 66% with the old method.

In a Nutshell
The paper says: "We don't need to show robots every second of a video to teach them. If we skip the boring parts and focus on the moments where the robot actually has to do something, the robot learns faster and does a better job."

It's the difference between reading a whole encyclopedia to learn how to tie a shoe versus just watching a 30-second video of someone tying a shoe. FRAMESKIP helps the robot find the "30-second video" inside the "encyclopedia."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →