← Latest papers
🤖 AI

OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents

OmniMapBench is a new benchmark comprising over 2,000 manually annotated question-answer pairs across diverse map documents, designed to evaluate and advance visual-centric reasoning in Large Vision-Language Models by introducing the Visual Dependency Index to measure irreducible visual grounding.

Original authors: Yang Chen, Yunwen Li, Yufan Shen, Minghao Liu, Tianyu Zheng, Bin Fu, Qunshu Lin, Zhi Yu, Botian Shi

Published 2026-07-13
📖 4 min read☕ Coffee break read

Original authors: Yang Chen, Yunwen Li, Yufan Shen, Minghao Liu, Tianyu Zheng, Bin Fu, Qunshu Lin, Zhi Yu, Botian Shi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're trying to teach a super-smart robot how to read a map. You might think, "No problem! Just have the robot read the text on the map, like a book." But here's the twist: maps are sneaky. They hide their secrets in lines, colors, and shapes that you can't just "read" with words.

This paper introduces a new challenge called OmniMapBench, which is basically a giant, tricky obstacle course made entirely of maps. The authors built this course to see if robots can actually see and think about what they see, or if they're just cheating by reading the text instead.

The Great "Text-Trap"

The paper argues that many previous tests for AI were like giving a student a math problem where the answer was hidden in the question's wording. If the robot could turn the picture into a list of words (like "70% of smokers feel..."), it could solve the puzzle without ever really looking at the image.

The authors say: "Stop cheating!" They argue that many existing tests are too easy because the pictures can be turned into text too easily. They explicitly rule out the idea that reading text descriptions is enough to solve map problems. In fact, they show that for some maps, if you try to describe the whole picture in words, you miss the crucial details needed to answer the question.

The New Challenge: A Map Jungle

To fix this, the team created OmniMapBench. Think of it as a massive library containing 1,603 different maps. These aren't just boring subway maps; they are a wild mix of:

  • Indoor navigation (finding a soda machine in a mall).
  • Topography (reading mountain heights and river paths).
  • Fantasy game maps (finding a building with an eye symbol).
  • Historical explorer routes (tracing a ship around Africa).

They hand-crafted 2,096 questions for these maps. Some are easy (just spotting a color), while others are super hard (requiring the robot to take three steps to figure out a route).

The "Visual Dependency" Test

How do you know if a robot is actually looking at the map or just guessing? The authors invented a clever test called the Visual Dependency Index (VDI).

Imagine you have a robot.

  1. Test A: You show it the map and ask, "How many houses are on the right?" It answers.
  2. Test B: You take the map away. Instead, you give the robot a long, boring paragraph describing the map in words (like "a green field with four clouds..."). Then you ask the same question.

If the robot gets Test B wrong, that's good! It means it needed the picture to solve the problem. The VDI measures exactly how much the robot's score drops when you take the picture away.

  • High VDI: The robot failed without the picture. It was actually looking!
  • Low VDI: The robot got it right even with just words. It was just reading the text.

The paper measured this and found that OmniMapBench has a much higher VDI than other tests. Even with a huge amount of text allowed to describe the image, the robots still struggled without the actual picture. This proves the test is fair and actually checks for visual skills.

The Results: Robots Are Still Learning

The authors put 25 of the smartest AI models (both free and paid ones) through this obstacle course.

  • The Best Score: The top robot, Gemini-3.1-Pro, got 75.03% correct.
  • The Gap: That might sound high, but in the world of AI, it means there's still a lot of room for improvement. The paper notes that even the best models struggle with the hardest, multi-step reasoning tasks.
  • The "Blind" Test: When they took away the pictures entirely and just gave the robots the questions and multiple-choice answers, the scores crashed from an average of 58.87% down to 23.35%. This proves the robots can't just guess based on the words; they really do need the map.

What's Next?

The paper doesn't claim this problem is solved. Instead, it suggests that current AI models are getting better at reading text but are still having a tough time with complex, visual reasoning. The authors hope that by using this new benchmark, researchers will build robots that can truly "see" the world, not just read the labels on it.

So, if you're a robot trying to find your way through a maze, don't just read the signs—look at the walls! The paper shows that right now, even the smartest robots are still learning how to do that.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →