LH-AVLN: A Benchmark for Long-Horizon Audio-Visual-Language Navigation
This paper introduces LH-AVLN, a new benchmark for long-horizon audio-visual-language navigation that challenges agents with multi-goal missions in acoustically rich indoor environments, and proposes PAG-Nav, a training-free agent that effectively leverages binaural audio for search while relying on visual-semantic verification for goal completion.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to be the ultimate house detective. Usually, when we teach robots to move around a house, we give them a simple mission: "Go find the red ball." They look around with their camera eyes, maybe remember where they've been, and walk until they spot it. This field of science is called "embodied navigation," where a robot learns to move through the real world to solve a task. But there's a catch: most of these training games are completely silent. The robot can only see what is directly in front of its lens. If the ball is behind a closed door or hidden in a dark corner, the robot is stuck, blind to anything outside its current view.
Now, imagine giving that robot a superpower: ears. In the real world, sound doesn't just travel in a straight line; it bounces around corners and seeps through cracks. A robot with good ears could hear a dog barking from the next room or a faucet dripping behind a wall, giving it a hint about where to go even when it can't see the object yet. This paper asks a big, tricky question: What happens if we combine these two superpowers—seeing and hearing—but make the mission much harder? Instead of finding just one thing, what if the robot has to find a whole list of different items, some described by words, some by pictures, and some just by their name, all while the sounds in the house keep changing as it solves the puzzle?
The researchers behind this study, led by Rufeng Chen and colleagues from the Hong Kong University of Science and Technology and Jilin University, have built a new challenge called LH-AVLN. Think of it as a high-stakes, multi-level video game for robots. In this game, the robot is dropped into a 3D house and given a "global mission" containing two to four different goals. These goals are tricky: one might be "find the sofa" (a category), another might be "find the light gray sofa next to the sheer curtains" (a description), and a third might be "find this specific picture of a bed" (an image reference).
The real twist, however, is the sound. In this game, every item the robot is looking for makes a noise. But here is the catch: the sounds are not fixed targets. As the robot finds and completes one goal, that item's noise stops. Meanwhile, the items it hasn't found yet keep making noise, taking turns to be the loudest. This creates a confusing situation. If the robot is looking for a sofa, but it hears a toilet flushing (because it hasn't found the toilet yet), the sound is a distraction. The robot has to figure out: "Is that sound helping me find my current target, or is it a red herring leading me to a different, unfinished task?"
To test this, the authors created a massive dataset with over 150,000 episodes where robots have to navigate these noisy, multi-goal missions. They found that existing robots, even the smart ones that are great at following instructions or remembering visual maps, struggle terribly. When the sound gets confusing or the mission gets long, these robots get lost or give up. They can't tell the difference between a helpful clue and a distracting noise.
To show how hard this is, the team built their own robot agent called PAG-Nav. This robot doesn't use complex training to learn; instead, it uses a clever strategy. It keeps a mental map of the whole house, tracking where it has been, what it has seen, and where the sounds are coming from. It uses the sound to guide its search when it can't see anything, but it refuses to say "I found it!" until it gets a good, clear look at the object to make sure it's the right one. Even with this smart strategy, the robot only succeeds in about 2% to 3% of the ordered missions and 3% to 5% of the unordered ones.
The main finding of this paper is that adding sound to long, complex missions makes things significantly harder, not easier, for current technology. The authors suggest that while sound is a powerful tool for finding things hidden from view, it becomes a nightmare if the robot can't figure out which sound belongs to which task. They argue that the future of robot navigation isn't just about seeing better or hearing better, but about learning to ignore the noise and focus on the right clue at the right time. The paper concludes that we are far from solving this puzzle, leaving plenty of room for future inventors to build robots that can truly navigate a noisy, multi-task world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.