← Latest papers
🤖 machine learning

VRIQ: Benchmarking and Analyzing Visual-Reasoning IQ of VLMs

The VRIQ benchmark reveals that current Vision Language Models perform poorly on visual reasoning tasks, primarily due to perception limitations rather than reasoning deficits, with accuracy ranging from 28% on abstract puzzles to 45% on natural images.

Original authors: Tina Khezresmaeilzadeh, Jike Zhong, Konstantinos Psounis

Published 2026-02-06
📖 4 min read☕ Coffee break read

Original authors: Tina Khezresmaeilzadeh, Jike Zhong, Konstantinos Psounis

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a team of incredibly smart robots that can read books and look at pictures. You want to know if they are truly "smart" or if they are just really good at guessing. To find out, the authors of this paper built a special test called VRIQ (Visual Reasoning IQ).

Think of VRIQ as a giant, digital "puzzle box" designed to trick these robots. The box contains two types of puzzles:

  1. Abstract Puzzles: These are like classic IQ test cards with weird shapes, lines, and patterns. There are no real-world objects here, just symbols.
  2. Natural Puzzles: These use real photos of everyday things, like apples, cars, or people, but ask the same tricky logic questions.

The Big Surprise: The Robots Are "Blind"

When the researchers tested the smartest AI models available (including the famous ones from OpenAI and Google), they found something shocking.

  • On Abstract Puzzles: The robots performed almost as badly as if they were just guessing randomly. Their average score was around 28% (where 25% is pure luck). It's like a student who has memorized the dictionary but can't read a single sentence.
  • On Natural Puzzles: They did a little better (around 45%), but they were still far from perfect.

The paper reveals that the problem isn't that the robots are bad at thinking (reasoning). The problem is that they are bad at seeing (perception).

The Detective Work: Why Did They Fail?

To figure out exactly where the robots were failing, the researchers acted like detectives. They didn't just ask, "What's the answer?" They broke every puzzle down into tiny steps using "probe questions."

Imagine a robot failing a puzzle where it has to find the odd shape out. The researchers asked:

  • Perception Probe: "How many red circles do you see?"
  • Reasoning Probe: "If there are 3 red circles and the rule is 'remove one,' what is left?"

The Results of the Investigation:

  • 56% of the time: The robot failed because it couldn't see the basics (e.g., it couldn't count the circles or tell which way a shape was pointing).
  • 43% of the time: The robot couldn't see the basics and couldn't figure out the logic.
  • Only 1% of the time: The robot saw everything perfectly but just got the logic wrong.

The Metaphor: It's like giving a chef a recipe for a cake. The chef (the AI) knows the logic of baking perfectly. But if the chef is blindfolded and can't see that there are only 2 eggs instead of 4, the cake will fail. The failure wasn't the cooking logic; it was the inability to see the ingredients.

The "Tool" Twist

The researchers tried one more thing. They gave one specific AI model (OpenAI's o3) a set of digital tools, like a magnifying glass, a pair of scissors (to crop images), and a calculator. This model was allowed to "think with images" by actively manipulating the picture before answering.

The Result: This tool-using robot soared. It jumped from a 28% score to over 50% on abstract puzzles and even 80-100% on natural puzzles. This proves that if you give the robot better "eyes" (tools to see clearly), its "brain" (reasoning) works much better.

The Bottom Line

The paper concludes that current AI models are not failing because they lack intelligence or logic. They are failing because they struggle to accurately perceive the visual world. They can't reliably count objects, understand 3D depth, or tell which way something is rotated.

To make AI truly smart at visual reasoning, we don't just need bigger brains; we need to give them better eyes.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →