← Latest papers
💻 computer science

Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control

The paper introduces HALO, a visuomotor policy that leverages attention-based memory retrieval distilled from vision-language model priors and sparse attention mechanisms to enable reliable long-horizon robot control in partially observable environments by mitigating spurious correlations and error accumulation.

Original authors: Rutav Shah, Yisu Li, Femi Bello, Yuke Zhu, Roberto Martín-Martín

Published 2026-06-25
📖 5 min read🧠 Deep dive

Original authors: Rutav Shah, Yisu Li, Femi Bello, Yuke Zhu, Roberto Martín-Martín

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to cook a complex meal in a kitchen. The problem is that the kitchen is huge, and the robot can only see what is directly in front of its eyes right now. It can't see the whole room, and it can't remember what happened five minutes ago unless you give it a special "memory."

This paper introduces a new way to give robots that memory, called HALO. Think of HALO as a smart, super-organized librarian for a robot's brain. Here is how it works, broken down into simple concepts:

The Problem: The Robot with Amnesia

Most robots today are like people with short-term memory loss. If you ask them to "put the bread in the microwave and then close the door," they might do the first part, but if you turn off the lights or move the bread out of sight, they forget what they were doing.

To fix this, scientists have tried to give robots a "notebook" where they write down everything they see. But there are two big problems with this approach:

  1. The "Wrong Clue" Problem: If the robot reads its whole notebook, it might get distracted by useless details. For example, if the robot sees a cat in the kitchen in the past, it might think, "Oh, the cat is here, so I should stop cooking!" even though the cat has nothing to do with the bread. This is called a "spurious correlation"—the robot connects two things that aren't actually related.
  2. The "Snowball Effect": If the robot makes a tiny mistake (like grabbing the bread slightly too hard), that mistake changes what it sees next. If it keeps reading its whole notebook while making mistakes, those errors pile up like a snowball rolling down a hill, until the robot is completely lost and fails the task.

The Solution: HALO (The Smart Librarian)

The authors created HALO to solve these two problems. It uses two main tricks:

Trick 1: The "Quiz Master" (Using AI to Teach Memory)

Imagine you are trying to learn a new language. You could just read a dictionary, but you learn faster if someone quizzes you on specific words you need to know.

HALO does this by using a "Quiz Master" (which is a powerful AI called a Vision-Language Model). Before the robot tries to cook, the Quiz Master looks at the robot's past experiences and asks it questions like:

  • "How many slices of bread were on the counter?"
  • "When did the stove get turned on?"
  • "Where was the olive oil placed?"

The robot has to answer these questions correctly to "pass the test." To answer, the robot must dig into its memory and find the specific, useful facts. This teaches the robot's memory system to ignore the useless stuff (like the cat) and focus only on the clues that actually help it cook.

Trick 2: The "Spotlight" (Ignoring the Noise)

Even with the Quiz Master, reading a whole 8-minute video of the robot's past can be overwhelming and confusing. It's like trying to find a specific sentence in a 500-page book by reading every single word.

HALO uses a "Spotlight" technique. Instead of reading the whole book, the robot uses a special search tool to find only the top 5 or 10 most important sentences from its past. It ignores everything else. This prevents the "Snowball Effect" because the robot isn't getting confused by old, noisy, or irrelevant information. It only looks at the specific moments that matter for the task it is doing right now.

The Results: How Well Did It Work?

The researchers tested HALO in computer simulations and with real robots in a lab. They gave the robots tasks that required remembering things from up to 8 minutes ago, such as:

  • Finding an object they saw earlier but couldn't see now.
  • Counting how many items were put in a drawer.
  • Waiting for a stove to heat up for a specific amount of time.

The findings were:

  • HALO was much better than previous methods. It succeeded at tasks about 41% of the time, while other methods (like robots that just read text summaries or robots with hand-written rules) only succeeded about 18% to 29% of the time.
  • The "Quiz Master" (VQA) helped the robot find the right clues, improving success by 10%.
  • The "Spotlight" (Top-k attention) stopped the robot from getting confused by its own mistakes, improving success by another 9%.

The Bottom Line

HALO teaches robots to be better at remembering by combining two things:

  1. Asking the right questions to learn what information is actually important.
  2. Looking at only the most important memories to avoid getting overwhelmed by noise.

This allows robots to handle long, complicated tasks in messy, real-world environments without getting lost or confused by their own history.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →