← Latest papers
🤖 AI

LifelongCrossNav: Persistent 3D Semantic Memory for Cross-Floor Multi-Object Navigation

LifelongCrossNav is a framework that enables sequential multi-object navigation across multiple floors in unknown indoor environments by maintaining a persistent shared sparse 3D semantic voxel memory and integrating specialized stair traversal capabilities, outperforming existing planar baselines on the newly introduced HM3D-MFMON benchmark.

Original authors: Zehui Li, Zihao Sun, Jiawei Xu, Zheqi He, Xiaoqiang Zhang, Jing-Shu Zheng, Lu Liu, Dahui Gao, Xiuwan Chen

Published 2026-08-10
📖 5 min read🧠 Deep dive

Original authors: Zehui Li, Zihao Sun, Jiawei Xu, Zheqi He, Xiaoqiang Zhang, Jing-Shu Zheng, Lu Liu, Dahui Gao, Xiuwan Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a robot that can walk through your house, find the remote, then the cat, and finally the toaster, all without getting lost or bumping into walls. This is the dream of "embodied AI," where computers don't just look at pictures but actually move through the world to get things done. For a long time, these robots were like people with very short memories and a fear of stairs. They could find one object on a single floor, but if you asked them to find three things in a row, or if the next item was upstairs, they would often get confused, forget where they had already looked, or simply refuse to climb the stairs. They treated the world like a flat map, flattening the second floor onto the first, which is great for a video game but terrible for a real house with a staircase.

The big question scientists have been asking is: How do we give a robot a "lifelong" memory that understands not just what things look like, but also how the floors connect? If a robot finds a chair in the living room, it should remember that chair even after it goes to the kitchen to find a spoon. And if the spoon is in the bedroom upstairs, the robot needs to know that stairs exist and how to use them, rather than just walking in circles on the first floor. This isn't just about being polite to your robot; it's about making machines that can actually help in complex, multi-story homes, hospitals, or offices where the action happens on multiple levels.

Enter LifelongCrossNav, a new framework designed to be the ultimate house-keeping robot with a supercharged memory. Think of it as giving the robot a mental 3D model of the entire house, built out of tiny, invisible blocks (voxels) that stack up to form walls, floors, and even the tricky geometry of a staircase. Unlike older robots that tried to flatten the world into a 2D paper map, LifelongCrossNav builds a true 3D structure. It remembers not just where things are, but how they are supported. It knows that a floor is solid because it's supported by the ground, but a gap in the air isn't walkable. Crucially, it treats stairs not as a scary obstacle, but as a special kind of path that connects different levels.

The system works like a curious explorer with a notebook. As the robot moves, it constantly updates its 3D map, adding new details about the shape of the room and the texture of the walls. It uses a special "vision-language" memory, which is like a super-advanced library catalog. When the robot is told to "find the TV," it doesn't just look for a box; it searches its memory for the visual and textual clues of a TV it might have seen earlier. If it found a TV in the living room during a previous task, it remembers that spot. If the next task is to find a bed, and the bed is upstairs, the robot doesn't panic. It checks its 3D map, sees the stairs, and plans a route up.

To test if this actually works, the researchers created a new challenge called HM3D-MFMON. Imagine a video game level with 36 different multi-story houses. They set up 927 scenarios where the robot had to find three different objects in a specific order. In 288 of these scenarios, the robot had to go upstairs or downstairs to finish the job; there was no way to do it on one floor. This was the ultimate stress test for the robot's ability to handle vertical movement and long-term memory.

The results were clear. When compared to a standard robot that uses a flat, 2D map (the baseline), LifelongCrossNav was significantly better at completing the full sequence of tasks. The flat-map robot often got stuck or failed completely when a floor transition was required, essentially giving up because it couldn't "see" the stairs in its flattened world. LifelongCrossNav, however, successfully navigated these cross-floor challenges. It didn't just find the objects; it found them more efficiently. By remembering where it had already looked (using "History POIs" or Points of Interest), it avoided re-walking the same halls, saving time and energy.

The paper suggests that this approach is a major step forward. It proves that giving a robot a persistent, 3D understanding of its environment—where it remembers the geometry of stairs and the location of objects across different floors—makes it much more capable of handling real-world tasks. While the robot still makes mistakes (sometimes confusing a bed for a sofa, for instance), the ability to navigate multiple floors and remember past discoveries without rebuilding the map from scratch is a solid, measurable improvement over previous methods. It's not a perfect robot yet, but it's the first one that truly understands how to live in a multi-story house.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →