Scaling Short-Term Memory of Visuomotor Policies for Long-Horizon Tasks
This paper introduces PRISM, a transformer-based architecture with gated attention and hierarchical compression that effectively scales short-term memory in visuomotor policies for long-horizon tasks, alongside the ReMemBench benchmark for systematic evaluation, achieving significant performance improvements over existing baselines without relying on large-scale pretraining.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to cook a complex meal, but every time you turn your head to grab a spice, you instantly forget what you were just doing. You might turn the stove on, walk away to get a pan, and then forget to turn it off, or worse, forget that you already put the salt in. This is the problem many current robots face: they are incredibly smart at seeing what is in front of them right now, but they have almost no "short-term memory" of what happened a few seconds or minutes ago.
This paper introduces a new robot brain called PRISM and a new test called REMEMBENCH to solve this problem. Here is how it works, explained simply:
The Problem: Robots with "Goldfish Memories"
Most modern robots learn by watching humans do tasks (like a student copying a teacher). However, they usually only look at the current video frame. If a task requires remembering something that happened 30 seconds ago—like "I turned the stove on, so I need to wait 2 minutes before turning it off"—the robot fails because it has no memory of that past action.
If you try to fix this by just giving the robot a longer video history to look at, two bad things happen:
- The Noise Problem: The robot gets overwhelmed by too much information. It might start connecting random, useless details (like a distractor object in the background) to its decisions, leading to mistakes.
- The Heavy Lifting Problem: Processing a long video history requires massive amounts of computer power and memory, making the robot slow and expensive to run.
The Solution: PRISM (The Smart Librarian)
The authors built PRISM, a robot policy that acts like a smart librarian rather than a hoarder. Instead of keeping every single page of a book in its head, it uses two clever tricks:
The "Gated Attention" Filter (The Bouncer):
Imagine the robot's memory is a crowded room. PRISM has a bouncer at the door who decides who gets to stay and who gets kicked out. If the robot remembers "I picked up a red apple," but the current task is about a blue bowl, the bouncer blocks the red apple memory from distracting the robot. This stops the robot from making silly connections between irrelevant past events and current actions.The "Hierarchical" Summarizer (The News Anchor):
Instead of reading every single word of a long history, PRISM first reads small chunks of the story and writes a short summary for each. Then, it reads those summaries to understand the whole plot. This is like a news anchor summarizing the day's events rather than reading every raw police report. This makes the robot much faster and requires less computer memory, allowing it to remember up to two minutes of history.
The Test: REMEMBENCH (The Robot Memory Exam)
To prove their robot actually has a memory, the authors created REMEMBENCH. Think of this as a standardized "driver's license test" for robot memory. It doesn't just test if the robot can move its arm; it tests four specific types of short-term memory:
- Spatial Memory: "Where did I put that fruit?" (The robot must find an object it can't currently see).
- Prospective Memory: "I need to turn off the stove in 2 minutes." (The robot must remember a future action based on a past trigger).
- Object-Associative Memory: "I washed this apple, so I must put it back in this specific bowl." (The robot remembers which object goes with which container).
- Object-Set Memory: "I need to move all the breadsticks, but I must count them as I go." (The robot keeps a running tally of a group of items).
The Results: A Clear Winner
When they tested PRISM against other robots (including those with standard "long memory" or different types of AI brains):
- On the Memory Test: PRISM was the clear winner, beating the next best robot by a significant margin (5% to 12% better).
- On Standard Tasks: Even on tasks that don't explicitly test memory, PRISM did better. Why? Because having a memory helps the robot understand context. For example, knowing if it "just picked up" a cup or is "about to put it down" helps it avoid confusion.
- Real World: When they tried it on a real physical robot, the robot without memory failed completely (0% success), while PRISM succeeded 30% of the time.
The Bottom Line
The paper claims that by using a "bouncer" to filter out noise and a "summarizer" to save space, robots can finally remember what happened a few minutes ago. This allows them to handle long, complex household chores that require a sequence of steps, rather than just reacting to the immediate moment. The authors provide the robot brain (PRISM) and the test (REMEMBENCH) to help other researchers build better, more memory-capable robots.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.