← Latest papers
💻 computer science

From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning

This paper introduces the Seeing-to-Experiencing (S2E) framework, which enhances navigation foundation models by combining offline pretraining on web-scale videos with reinforcement learning in simulation to improve interactive reasoning and safety, supported by a new evaluation benchmark called NavBench-GS.

Original authors: Honglin He, Yukai Ma, Brad Squicciarini, Wayne Wu, Bolei Zhou

Published 2026-06-12
📖 5 min read🧠 Deep dive

Original authors: Honglin He, Yukai Ma, Brad Squicciarini, Wayne Wu, Bolei Zhou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Tourist" vs. The "Local"

Imagine you want to teach a robot how to walk through a busy city.

The Old Way (Just "Seeing"):
Currently, most navigation robots are trained like a tourist who has watched thousands of hours of YouTube videos about walking in cities. They have seen every type of street, crowd, and obstacle on screen. This is called Offline Pre-training.

  • The Flaw: Watching a video of someone dodging a skateboarder is very different from actually dodging one. The robot knows what a collision looks like, but it doesn't understand the physics of falling or the split-second timing needed to stop. If the robot tries to walk in the real world, it might freeze or crash because it has never actually "felt" the consequences of a wrong turn. It lacks causality (cause and effect).

The New Way (The "S2E" Framework):
The authors propose a new method called Seeing-to-Experiencing (S2E). Think of this as taking that tourist and sending them to a highly realistic, safe video game simulator where they can actually walk, trip, and learn from their mistakes without getting hurt.

The goal is to combine the broad knowledge of the video-watching tourist with the street smarts of someone who has actually practiced in the real world.


How It Works: The Two-Step Recipe

The S2E framework uses two main "tricks" to make this work efficiently.

1. The "Anchor" System (During the Video Phase)

The Challenge: When a robot watches a video, it sees that sometimes people walk straight, sometimes they turn left to avoid a dog, and sometimes they stop for a red light. All these actions are valid for the same situation. If the robot tries to guess just one average path, it will fail.

The Solution: The authors use something called Anchor-Guided Distribution Matching.

  • The Analogy: Imagine the robot is a chef trying to guess a recipe. Instead of guessing a single flavor, they place several "Anchors" (like specific flavor profiles: "Spicy," "Sweet," "Savory") on the table.
  • How it helps: The robot learns that for a specific situation, the answer might be "Spicy" (go straight) OR "Savory" (turn left). By using these anchors, the robot learns to predict a menu of possible good actions rather than just one average guess. This makes the robot much more flexible and ready for different scenarios.

2. The "Residual" Brain (During the Practice Phase)

The Challenge: Once the robot starts practicing in the simulator, we want it to learn new tricks (like dodging a moving pedestrian). But if we let the robot relearn everything from scratch, it might forget how to walk straight or recognize a sidewalk. This is called "catastrophic forgetting." Also, the simulator looks slightly different from the real world (different lighting, textures), which can confuse the robot.

The Solution: They use a Residual-Attention Module.

  • The Analogy: Imagine the robot has a "Base Brain" (the part trained on videos) that is excellent at recognizing streets. When it enters the simulator, we don't replace the Base Brain. Instead, we attach a small, lightweight "Add-on Brain" (the Residual Module) on top of it.
  • How it helps: The Base Brain stays frozen and keeps its general knowledge safe. The Add-on Brain is the only part that learns. It focuses only on the new, interactive stuff: "Oh, that person is moving, I need to slow down."
  • The Benefit: This is like adding a new app to your phone without deleting your contacts. The robot gains new reactive skills (dodging, stopping) without losing its old general knowledge (recognizing a road).

The Test Drive: NavBench-GS

To prove this works, the authors built a new testing ground called NavBench-GS.

  • What is it? Instead of testing on flat 2D images (like looking at a photo of a street), they built a 3D world that looks exactly like the real world but has real physics. You can throw a ball at the robot, and it will react.
  • The Results:
    • Robots trained only on videos (the "Tourists") crashed often when faced with moving people or obstacles.
    • Robots trained with the S2E method (the "Locals") were much better. They reached their destination more often and crashed far less.
    • The Efficiency Surprise: The S2E robot learned faster with less data. A robot trained on just 100 hours of data + simulator practice performed better than other robots trained on over 2,000 hours of video data. This proves that interactive practice is more valuable than just watching more videos.

Real-World Proof

The team didn't just stop at simulations. They put their robot on two real machines:

  1. A wheeled robot (like a delivery bot).
  2. A quadruped robot (a dog-like robot).

They tested these robots in real city streets with real obstacles and people. The S2E robots successfully navigated these complex environments without prior specific training on those exact streets (Zero-Shot Generalization). They could keep to the sidewalk and avoid pedestrians, proving that the skills learned in the simulator transferred perfectly to the real world.

Summary

The paper argues that to make robots truly smart navigators, we can't just feed them more videos. We must let them practice. By combining a stable "video-based" foundation with a "practice-based" learning module that doesn't overwrite their original knowledge, robots can learn to navigate safely and efficiently in the chaotic, real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →