← Latest papers
💻 computer science

SLVMBench: Skill Learning from Video Memory

This paper introduces SLVMBench, the first benchmark designed to evaluate video large language models' ability to learn procedural skills from long video streams containing embedded tutorials and apply them to real-time tasks, revealing that current models struggle significantly with this capability due to limitations in long-context memory and skill transfer.

Original authors: Yudong Yang, Guangzhi Sun, Yixuan Li, Chao Zhang

Published 2026-07-14
📖 5 min read🧠 Deep dive

Original authors: Yudong Yang, Guangzhi Sun, Yixuan Li, Chao Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're trying to learn how to fix a leaky faucet. You find a perfect, 10-minute tutorial video on YouTube. You watch it, nodding along, feeling like you've got the whole thing down. Then, you close the tab and go do something else for two hours. Maybe you watch a bunch of random cat videos, a cooking show, and a documentary about deep-sea fish.

Now, you go back to the sink. The faucet is dripping. You pick up your wrench, but suddenly, your brain goes blank. Wait, did I loosen the nut first or the screw? Was it a wrench or a plier? You can't remember the tutorial because the two hours of "cat video noise" drowned it out.

This is exactly the problem SLVMBench (Skill Learning from Video Memory) is trying to solve, and it's the first time anyone has built a test to see if AI can do what humans do: learn a skill from a video, forget it for a while, and then remember it when they actually need to use it.

The Big Test: The "Two-Hour Distraction" Challenge

The researchers created a super-tough game for AI models. Here's how it works:

  1. The Tutorial: The AI watches a video teaching it a specific skill (like how to use a power tool or fix a gadget).
  2. The Noise: Immediately after, the AI is forced to watch a stream of 2 hours of completely irrelevant, boring, or random videos. This is the "distractor" phase.
  3. The Test: The AI is then shown a new video of someone doing the task, but the video stops right before a crucial step. The AI has to guess what happens next, using only the memory of that tutorial from two hours ago.

Think of it like a memory game where you have to remember a secret recipe, but someone keeps shouting random numbers at you for two hours before asking, "What's the first ingredient?"

The Shocking Result: AI Has a "Memory Cliff"

The paper tested the smartest AI models out there, including big names like Gemini, GPT-4o, and GPT-5.2. The results were a bit of a reality check.

When the AI was asked to solve the problem immediately after watching the tutorial (with no distraction), it did pretty well. It was like asking you to fix the faucet right after the tutorial video ended. You'd probably get it right.

But the moment they added the 2-hour distraction, the AI's performance took a nosedive.

  • The "Cliff": The paper describes a "memory cliff." For some of the top models, their ability to remember the skill dropped so low that they were barely better than if they had never seen the tutorial at all.
  • The Numbers: One of the best models, Gemini 3.1 Pro, saw its score drop from a solid 74.92% (with immediate help) down to 59.45% after the long distraction. That's a big gap. But for other models like GPT-5.2, the drop was even worse: from 57.94% down to 44.89%. In fact, for some models, the long wait made their performance almost identical to just guessing without any help at all.

The authors suggest that while these AIs are great at watching a video and answering questions while it's playing, they are terrible at holding onto that information when it's buried under hours of other stuff.

What the Paper Rules Out (And What It Doesn't)

The paper is very clear about what this test is not about.

  • It's not about "Common Sense": The questions were designed so you couldn't guess the answer just by being smart. You had to remember the specific tutorial. If an AI got it right without the tutorial, the researchers threw that question out.
  • It's not about "Passive Watching": Previous tests just asked, "What happened in this 1-hour video?" This test asks, "What do you do next based on a video you saw hours ago?" It's about doing, not just remembering facts.
  • It's not a "Win" for Streaming AI: Some people thought that models designed to watch long streams (like "streaming" models) would be better at this. The paper shows that while they are slightly better, they still struggle mightily. One model, PEMF, saw its accuracy plummet from 64.88% to 39.36% when the distraction was added. This suggests that just "watching longer" isn't the magic fix.

How Sure Are We?

The researchers didn't just guess; they built a massive, carefully crafted dataset.

  • The Data: They collected 1,220 videos and created 2,261 questions.
  • Human Proof: Real humans spent over 750 hours checking every single video and question. They made sure the tutorial actually taught the skill, that the "distractor" videos were truly irrelevant, and that the questions were timed down to the sub-second level.
  • The Confidence: The paper reports these results with 95% confidence intervals, meaning they are statistically sure these drops in performance are real and not just a fluke.

The Takeaway

The paper concludes that current AI models are like students who can ace a test if they study right before the exam, but if they have to study two hours before and then sit through a boring lecture, they forget everything.

The authors suggest that for AI to truly act like a helpful robot in the real world—learning a skill from a video and using it days later—we need to figure out how to stop these "memory cliffs." Until then, even the smartest AI models are still struggling to keep their skills sharp when the noise gets loud.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →