← Latest papers
💻 computer science

NaviTrace: Evaluating Embodied Navigation of Vision-Language Models

This paper introduces NaviTrace, a high-quality Visual Question Answering benchmark featuring over 3000 expert navigation traces across diverse scenarios and embodiment types, which reveals that current vision-language models still lag behind human performance in spatial grounding and goal localization when evaluated with a novel semantic-aware trace score.

Original authors: Tim Windecker, Manthan Patel, Moritz Reuss, Richard Schwarzkopf, Cesar Cadena, Rudolf Lioutikov, Marco Hutter, Jonas Frey

Published 2026-03-10
📖 5 min read🧠 Deep dive

Original authors: Tim Windecker, Manthan Patel, Moritz Reuss, Richard Schwarzkopf, Cesar Cadena, Rudolf Lioutikov, Marco Hutter, Jonas Frey

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a brilliant but inexperienced tourist how to navigate a new city. You show them a photo of a street and say, "Walk to that red car." A human tourist would look at the photo, spot the car, notice the stairs, see the "Do Not Walk" sign, and draw a path on the photo showing exactly where to step.

This paper introduces NaviTrace, a new "test" designed to see if modern AI robots (specifically Vision-Language Models) can do the same thing.

Here is a breakdown of the paper using simple analogies:

1. The Problem: The "Robot Tourist" is Getting Lost

Robots are getting smarter at understanding language and seeing pictures. But when you ask a robot to "go downstairs" or "follow the road," we don't really know how good it is at figuring out the path.

  • The Old Way: To test robots, researchers used to send them out into the real world. This is like testing a driver by having them drive across the country every time you want to check their skills. It's expensive, slow, and hard to repeat.
  • The Simulation Way: Others use video game simulations. This is like testing a driver in a flight simulator. It's cheap and repeatable, but the "roads" in the game are often too simple, missing real-world details like uneven pavement or social rules (like waiting for a pedestrian).

2. The Solution: A "Paper-and-Pencil" Navigation Exam

The authors created NaviTrace, which is like a standardized exam for robot navigation.

  • The Setup: Instead of sending a robot out, they show an AI a single photo of a real-world scene (like a busy street or a park) and give it a command (e.g., "Go to the end of the road").
  • The Twist: They ask the AI to draw a 2D line (a "trace") directly on the photo showing where a specific type of traveler should go.
  • The Travelers: They test four different "embodiments" (types of travelers):
    1. Human: Can walk anywhere but can't climb tall walls.
    2. Legged Robot: Like a dog on four legs; can handle rough ground but is short.
    3. Wheeled Robot: Like a delivery bot or wheelchair; needs smooth paths and ramps.
    4. Bicycle: Needs to stay on roads or bike lanes; hates stairs.

3. The Grading System: The "Smart Ruler"

How do you grade a drawing? You can't just measure the distance from start to finish. The authors invented a special scoring system that acts like a smart ruler:

  • Shape Match: Does the drawn line look like the path a human expert would draw? (They use a math trick called "Dynamic Time Warping" to compare the shapes).
  • Goal Check: Did the line actually end up at the destination?
  • The "Don't Go There" Penalty: This is the clever part. The system knows what things are in the photo (e.g., grass, stairs, a busy road). If the AI draws a line for a wheelchair going up a staircase, the system gives it a heavy penalty. If it draws a line for a human walking on a bike lane, it gets a smaller penalty.

4. The Results: The AI is Smart, But Not "Grounded"

The authors tested top-tier AI models (like Gemini, GPT-5, and others) against human experts.

  • The Gap: Humans scored around 75/100. The best AI models scored around 34/100. The AI is doing better than random guessing, but it is still far behind a human.
  • The Main Mistake: The biggest problem isn't that the AI doesn't know how to walk; it's that it can't find the destination.
    • Analogy: If you tell the AI "Go to the red car," it often draws a path that looks reasonable but ends up at the wrong car, or it draws a path that goes off the map entirely.
    • When the researchers forced the AI to just guess the destination and then drew a straight line to it, the score didn't improve much. This means the AI is struggling to understand where things are in the picture, not just how to move.
  • The "Reasoning" Trap: Some AI models (like o3) wrote out perfect logical steps in text: "I need to go left, avoid the stairs, and head to the car." However, when they actually drew the line, they ignored their own text and drew a path that went into a wall. The AI's "brain" (text) and its "eyes" (drawing) aren't talking to each other properly.

5. Why This Matters

The paper argues that before we can trust robots to deliver packages, help the elderly, or search for survivors, we need a way to test them that is fair, cheap, and covers real-world complexity. NaviTrace provides that test. It shows us that while AI is great at chatting and recognizing objects, it still struggles to "see" the world in a way that allows it to move safely and logically through it.

In short: The paper built a navigation test where robots have to draw a map on a photo. The results show that while the robots are getting better, they are still terrible at finding their way and often ignore the physical rules of the road (like stairs for wheelchairs).

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →