← Latest papers
💻 computer science

Do Open-Loop Metrics Predict Closed-Loop Driving? A Cross-Benchmark Correlation Study of NAVSIM and Bench2Drive

This study analyzes the correlation between NAVSIM open-loop metrics and Bench2Drive closed-loop performance across state-of-the-art methods, revealing that while aggregate NAVSIM scores show strong but non-monotonic correlation with real-world driving success, Ego Progress is the most predictive single metric and a simplified three-metric formula can effectively replace the full five-metric score for ranking purposes.

Original authors: Yiru Wang, Anqing Jiang, Shuo Wang, Yuwen Heng, Hai Yang, Yang Chen, Hao Sun

Published 2026-05-04
📖 5 min read🧠 Deep dive

Original authors: Yiru Wang, Anqing Jiang, Shuo Wang, Yuwen Heng, Hai Yang, Yang Chen, Hao Sun

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to drive a car. To see if the robot is good, you have two ways to test it:

  1. The "Paper Test" (Open-Loop): You give the robot a map and a set of instructions, and you ask it to draw a line on a piece of paper showing where it would go. You then compare that line to the perfect path. This is fast, cheap, and you can do it a thousand times in a second.
  2. The "Real Drive" (Closed-Loop): You actually put the robot in a car (or a super-realistic video game) and let it drive. The car reacts to traffic lights, other cars, and pedestrians. If the robot gets stuck or drives too slowly, it fails. This is the "gold standard," but it takes hours, costs a lot of money, and is hard to repeat exactly the same way.

The Big Question: Can the "Paper Test" reliably predict how the robot will do in the "Real Drive"?

For a long time, the answer was no. Old tests just measured how far off the robot's drawing was from the perfect line (like measuring the distance between two dots). The paper shows that even if a robot draws a nearly perfect line, it might still crash or get stuck in real life.

What This Paper Did

The authors looked at 8 different top-tier robot drivers. They checked how these robots did on the "Paper Test" (using a new, smarter test called NAVSIM) and compared it to how they actually did on the "Real Drive" (using a test called Bench2Drive).

Here is what they found, explained simply:

1. The "Safety vs. Speed" Trap

The new "Paper Test" (NAVSIM) is much better than the old ones. It checks for safety (did you hit anything?) and progress (did you move forward?).

  • The Problem: Some robots are too safe. They drive like a turtle, stopping at every little shadow to make sure they don't hit anything.
  • The Result: On the "Paper Test," these turtle-robots get high scores because they never hit anything. But in the "Real Drive," they get penalized for driving too slowly or getting stuck in traffic. They fail the real test because they are too cautious.
  • The Analogy: Imagine a student who answers every question on a test by saying "I don't know" to avoid getting a single point wrong. On a test that only counts wrong answers, they get an A. But on a test that requires you to actually finish the exam, they get an F because they didn't finish.

2. The Secret Ingredient: "Ego Progress"

The researchers broke down the "Paper Test" into its parts to see which one actually predicted real-world success.

  • They found that the most important number wasn't "Did you crash?" (Safety). It was "Did you move forward?" (Progress).
  • The Analogy: Think of a marathon. If you run perfectly without tripping (Safety) but you stop to tie your shoe every 10 seconds, you won't win the race. The robot needs to keep moving forward. The "Progress" score was the single best predictor of whether the robot would actually finish the drive successfully.

3. The "Snowball Effect"

Why do small mistakes in the "Paper Test" become huge failures in the "Real Drive"?

  • The Analogy: Imagine you are walking down a hallway. If you take one tiny step to the left, it doesn't matter. But if you keep taking tiny steps to the left for 10 minutes, you will eventually walk into a wall.
  • In the "Paper Test," the robot only plans 4 seconds ahead. A tiny hesitation might look like a small error. But in the "Real Drive," that tiny hesitation causes the robot to miss a green light, wait for a red light, fall behind schedule, and eventually get timed out. Small errors "snowball" into total failure.

4. A Simpler Formula

The "Paper Test" uses a complex formula with 5 different numbers to calculate a final score. The authors realized that two of those numbers (how comfortable the ride is and how close you are to a crash) were basically the same for all the top robots—they were all perfect at those things.

  • The Discovery: You can throw away the complex formula and just use 3 simple numbers: "Did you crash?", "Did you stay in your lane?", and "Did you move forward?"
  • The Result: This super-simple formula predicted the real-world results just as well as the complex one. It's like realizing you don't need a 50-page report to know if a car is good; you just need to know if it moves and doesn't crash.

The Bottom Line

The paper concludes that while we can't perfectly predict real driving just by looking at a drawing, we are getting much closer.

  • Don't just look at safety: A robot that is safe but doesn't move is a bad driver.
  • Watch the "Progress": The ability to keep moving forward is the best sign of a good driver.
  • Simplify: We don't need overly complex scoring systems to tell good drivers from bad ones; a simple focus on safety and movement works best right now.

The authors warn that as robots get better, we might need to tweak these rules again, but for now, this "Progress" metric is the key to knowing who will actually succeed on the road.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →