← Latest papers
💻 computer science

Look Up and Look Back: Hidden Attention and Latent Orientation in a Frozen Foundation Model for Panoramic SLAM

The paper presents HALO-SLAM, a novel monocular panoramic SLAM system that leverages hidden attention and latent orientation cues from a frozen foundation model to achieve IMU-free gravity alignment and robust loop closure, resulting in 100% success and state-of-the-art accuracy across five real-world benchmarks.

Original authors: Zhuang Xiong, Guohao Zhang, Chen Zhang, Zheyu Jiang, Yuchao Mei, Qingshan Xu, Wenbing Tao

Published 2026-08-04
📖 2 min read☕ Coffee break read

Original authors: Zhuang Xiong, Guohao Zhang, Chen Zhang, Zheyu Jiang, Yuchao Mei, Qingshan Xu, Wenbing Tao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to build a giant, 3D map of the world using only a single camera that can see everything around it at once, like a fish-eye lens on steroids. This is the challenge of "Panoramic SLAM" (Simultaneous Localization and Mapping). It's the technology that lets robots and self-driving cars know where they are and what the world looks like without GPS. The tricky part is that these cameras see the whole world in a distorted, flat rectangle (like a world map), and if the camera tilts even a little, the whole picture gets scrambled. For a long time, computers struggled to keep their bearings when the camera rolled or pitched, often getting lost or building maps that stretched and warped like melted wax. They needed a way to figure out which way was "up" and to recognize when they had returned to a place they had already visited, all while keeping the map consistent.

Now, meet HALO-SLAM, a new system that acts like a super-smart detective for these panoramic cameras. Instead of trying to learn everything from scratch, HALO-SLAM uses a "frozen" giant AI brain (a foundation model called PanoVGGT) that was already trained to understand 3D shapes. The clever trick is that HALO-SLAM doesn't just look at the final answer the AI gives; it peeks inside the AI's "thought process" (its hidden layers) to find secret clues. First, it reads the AI's internal tokens to guess which way is down, allowing it to straighten the camera's view without needing a physical gyroscope. Second, it uses the AI's attention mechanism—essentially seeing how much the AI "looks" at one image while thinking about another—to decide if two photos are truly of the same place or just look similar by accident. By combining these hidden clues with a rigorous three-step check, HALO-SLAM managed to successfully map 125 different real-world sequences, achieving a 100% success rate and reducing measurement errors by up to 88% compared to the best previous methods. It proves that by listening to the quiet whispers inside a pre-trained AI, we can build much more reliable maps for robots and virtual reality.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →