Retrospective Open-Vocabulary Memory for Long-Term Object Search
This paper introduces ECROM, a method for long-term object search that utilizes retrospective open-vocabulary memory to reason about detection opportunities and improve search efficiency, validated by a new benchmark showing significant performance gains over existing memory systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Robots that move through the world are getting better at seeing, but they are still terrible at remembering where things usually are. Imagine a robot tasked with finding a lost set of keys in a house it has visited many times. If the robot only remembers the very last time it saw the keys, it might look in the wrong place if the keys were moved yesterday. If it simply counts how many times it has seen the keys in a specific spot, it might be fooled by its own bad luck: perhaps it walked past the kitchen table a hundred times but never looked down at it, while it only glanced at the living room sofa once and happened to see the keys there. A simple count would make the robot think the sofa is the best place to look, even though the table is actually where the keys belong most of the time. This is the core problem of long-term object search: distinguishing between where an object actually lives and where the robot just happened to look.
To solve this, researchers at the National University of Singapore have developed a new way for robots to build a memory of their environment. They call their system ECROM, which stands for Exposure-Calibrated Retrospective Open-Vocabulary Memory. The team tested this system in a series of controlled simulations across ten different digital homes, where objects were moved around according to hidden patterns, and the robot's path was deliberately varied so that some surfaces were seen often and others rarely. They found that by carefully weighing what the robot saw against how much of a surface it actually had a chance to see, the robot could learn the true habits of objects. When asked to find items it had never been told to look for before, this new memory system helped the robot find the target much faster and more reliably than previous methods.
The challenge the researchers tackled is rooted in the messy reality of how robots move. In a real home, a robot does not scan every inch of every room with perfect clarity. It might walk through a kitchen while looking at the ceiling, or pass a dining table that is partially blocked by a chair. If a robot simply counts how many times it detected a "straw hat" on a table versus an island, it might get the answer wrong if it spent more time looking at the island. The hat might be on the table half the time, but if the robot only ever looked at the island, a simple count would suggest the island is the hat's home. The researchers realized that to learn where things usually occur, a robot must separate the evidence of an object's presence from the opportunity it had to see it. A missed detection is only meaningful if the robot was actually looking in the right direction; if the robot was looking away, the absence of a detection tells us nothing.
To test this idea, the team created a benchmark called LOOP-Bench, which consists of ten realistic home environments. In these simulations, they placed objects like cups, scissors, and hats on various surfaces according to hidden probability distributions. Some objects were always on the same table, while others moved between a few spots, and some rarely appeared at all. Crucially, the robot's path through these homes was generated independently of where the objects were placed. This meant the robot might walk past a table ten times without ever seeing the surface clearly, while it might only pass a kitchen island once but get a perfect view. This setup forced the memory system to figure out the true location of the objects without being tricked by the robot's own movement patterns.
The new system, ECROM, works by organizing the robot's history into a map of surfaces, or "supports," and recording exactly how much of each surface was visible during every pass. When a human later asks the robot to find a specific item, the system looks at all the past trips. It does not just count the detections; instead, it calculates a probability based on how much of the surface was exposed to the camera. If the robot saw a surface clearly and did not find the object, the system strongly lowers the belief that the object is there. But if the robot only caught a glimpse of a surface, or if the view was blocked, the system treats the lack of a detection as neutral, refusing to penalize that spot. This allows the robot to learn that an object is likely to be on a surface it has rarely seen, provided that surface is the only one where the object has ever been spotted when it was visible.
The results of the simulation were clear. When the researchers asked the robot to find objects it had never been explicitly told to search for, the new system outperformed every other method they tested. It improved the accuracy of its predictions by a significant margin and, more importantly, it found the objects much faster during active search. In the simulations, the robot using ECROM found the target in about 59 percent of its attempts, compared to roughly 53 percent for the next best method. More strikingly, in a specific case study shown in the paper, the distance the robot had to travel to find the object dropped from 17.8 meters to just 3.2 meters. This efficiency came from the robot knowing exactly which surfaces were worth checking first, rather than wandering aimlessly or getting stuck revisiting places it had already ruled out.
To ensure this was not just a digital trick, the team also tested the system on a physical robot, a Hello Robot Stretch 3, in a real office environment. They ran the same type of search task with real objects like a red cup, a pair of scissors, and a can of luncheon meat. The robot had to navigate a 500-square-meter space, moving between desks and shelves, to find these items. The results mirrored the simulation. The robot using the new memory system succeeded in finding the objects in 14 out of 15 trials, while the other methods succeeded in only 10 or 11. The physical robot also traveled significantly less distance, covering an average of 19.4 meters compared to nearly 29 meters for the other approaches. This proved that the logic of separating observation opportunity from object presence works even when the robot is dealing with the unpredictable lighting and clutter of the real world.
The researchers also explored what would happen if they removed the key parts of their system. When they forced the robot to assume it had seen every surface fully, even when it clearly hadn't, the robot's performance dropped. It began to treat missed detections as proof that an object was absent, leading it to ignore the correct locations. Similarly, when they simplified the system to only look at whether an object was seen or not, ignoring the strength of the visual evidence, the robot became less precise. These tests confirmed that the system's success relied on two specific things: a careful accounting of how much of a surface was actually visible, and a nuanced interpretation of how strong the visual evidence was.
This work suggests a path forward for robots that need to operate in dynamic, changing environments over long periods. Instead of just remembering the last time they saw something, or keeping a simple tally of sightings, these robots can learn the habits of the world around them. They can understand that just because they haven't seen a coffee mug on a counter recently, it doesn't mean the mug isn't there; it might just mean they haven't looked at that counter in a while. By building a memory that respects the limits of its own vision, a robot can become a more reliable partner in the home, capable of finding lost items even when the world around it is in constant flux. The researchers plan to release their code and the benchmark data to the public, allowing others to build upon this foundation of smarter, more observant machines.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.