← Latest papers
💻 computer science

NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation

NavWAM is a diffusion-transformer policy that unifies future visual prediction, goal-progress value estimation, and action generation into a shared latent sequence, enabling a robot to directly execute closed-loop navigation without relying on external planners or action search.

Original authors: Daichi Azuma, Taiki Miyanishi, Koya Sakamoto, Shuhei Kurita, Yaonan Zhu, Petr Khrapchenkov, Motoaki Kawanabe, Yusuke Iwasawa, Yutaka Matsuo

Published 2026-06-12
📖 5 min read🧠 Deep dive

Original authors: Daichi Azuma, Taiki Miyanishi, Koya Sakamoto, Shuhei Kurita, Yaonan Zhu, Petr Khrapchenkov, Motoaki Kawanabe, Yusuke Iwasawa, Yutaka Matsuo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to find a specific room in a house, but the robot can only see what's directly in front of its "eyes" (a camera). It doesn't have a map, and it can't see around corners. This is the challenge of goal-conditioned visual navigation.

Here is the story of how the authors, Daichi Azuma and his team, solved this problem with a new system called NavWAM.

The Problem: The "Crystal Ball" vs. The "Driver"

In the past, researchers tried to teach robots to navigate using two separate tools:

  1. The Crystal Ball (World Model): This tool is great at guessing what the robot will see if it moves forward. "If I turn left, I'll see a kitchen." "If I go straight, I'll hit a wall."
  2. The Driver (Planner): This tool looks at the Crystal Ball's guesses and decides, "Okay, the kitchen looks good, so I'll turn left."

The Flaw: The paper argues that keeping these two separate is inefficient. It's like having a passenger who screams out directions ("Turn left!") while the driver has to stop, think, and then turn. It creates a bottleneck. The "Crystal Ball" predicts the future, but it doesn't know how to drive the car to get there.

The Solution: NavWAM (The "Driver-Predictor")

The authors created NavWAM (Navigation World Action Model). Think of NavWAM as a super-driver who can also see the future.

Instead of having a separate "Crystal Ball" and a "Driver," NavWAM combines them into one brain. It doesn't just guess what it will see; it guesses what it will see, how close it is to the goal, and exactly what steering commands to give—all at the same time.

The Creative Analogy: The "Shared Canvas"

To understand how NavWAM works, imagine a painter working on a single canvas with nine different panels:

  • The Bottom Panels (What we know): These show the robot's current view, where it is right now, and the picture of the goal (e.g., a photo of a specific chair).
  • The Top Panels (What we predict): The robot paints these panels to show:
    1. What the robot will see in the future.
    2. How close it will be to the goal.
    3. The most important part: The actual steering commands (the "action chunk") needed to get there.

The robot learns by "denoising" this canvas. It starts with a messy, blurry version of the future and cleans it up until it has a clear picture of the path, the destination, and the steps to take. Because the "steering commands" are painted on the same canvas as the "future view," the robot learns that to see the goal, it must take these specific actions.

How It Beats the Competition

The paper tested NavWAM against two other types of robots:

  1. The Old Way (NWM): The robot predicts the future, then uses a complex, slow math search (called CEM) to figure out which action to take. It's like a chess player calculating 100 moves before making one.
  2. The Direct Way (OmniVLA): The robot just looks at the goal and guesses the next move without thinking about the future. It's like a driver who just reacts to the road without looking ahead.

The Results:

  • NavWAM vs. The Old Way: NavWAM was much faster and more accurate because it didn't need to stop and calculate a complex search. It just "knew" the move. It reached the goal in 79% of real-world tests, while the Old Way only succeeded 16% of the time.
  • NavWAM vs. The Direct Way: NavWAM performed just as well as the Direct Way (which uses a much larger, more powerful computer brain), but NavWAM had the added superpower of actually predicting what the future would look like.

The Real-World Test

The team didn't just test this in a computer simulation. They put NavWAM on a real robot named Diablo and sent it into four different indoor environments (an office, a storage room, a meeting room, and a hallway).

  • The Goal: Find a specific image captured in that room.
  • The Outcome: NavWAM successfully navigated to the goal in 19 out of 24 attempts.
  • The "Foresight" Check: Even though NavWAM didn't need to use its future predictions to move (it just used the steering commands), the paper shows that its predictions were surprisingly accurate. When the robot predicted "I will see a door in 4 seconds," it actually saw a door in 4 seconds.

The Big Takeaway

The paper concludes that for a robot to navigate well, it shouldn't just be a "predictor" or just a "driver." It needs to be both at the same time. By learning to predict the future and the actions required to create that future in one single step, the robot becomes much better at finding its way, even in places it has never seen before.

In short: NavWAM teaches the robot to "drive while looking ahead," rather than "look ahead, stop, calculate, and then drive."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →