← Latest papers
🤖 AI

A Study of Commonsense Reasoning over Visual Object Properties

This paper introduces the OPTICS benchmark and a systematic evaluation framework to assess commonsense reasoning over visual object properties in vision-language models, revealing that even state-of-the-art models significantly lag behind human performance, particularly in handling photographic images, counterfactuals, and complex physical or functional attributes.

Original authors: Abhishek Kolari, Mohammadhossein Khojasteh, Yifan Jiang, Floris den Hengst, Filip Ilievski

Published 2026-07-09
📖 5 min read🧠 Deep dive

Original authors: Abhishek Kolari, Mohammadhossein Khojasteh, Yifan Jiang, Floris den Hengst, Filip Ilievski

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to look at a picture and answer questions about it, like "How many cats are in this photo?" or "If we took away the red chairs, how many blue ones would be left?"

This paper is like a report card for the smartest robots (called Vision-Language Models) we have today. The researchers wanted to see if these robots truly "understand" what they see, or if they are just guessing based on patterns they've memorized.

Here is the story of their study, broken down into simple parts:

1. The Problem: The "Blindfolded Math" Test

Previous tests for these robots were a bit like a messy kitchen. They mixed up simple tasks (like spotting a red ball) with hard tasks (like figuring out if a ball is heavy enough to break a window). Because everything was mixed together, it was hard to tell if the robot was bad at seeing or bad at thinking.

Also, most tests used perfect, cartoon-like drawings. Real life is messy, with shadows, blurry edges, and things hiding behind other things. The researchers wanted to know: Can these robots handle the messy real world, or do they only work in a perfect cartoon world?

2. The Solution: The "OPTICS" Test Kitchen

To fix this, the team built a new, very organized test called OPTICS. Think of it as a three-course meal designed to test the robot's brain in three specific ways:

  • The Menu (Object Properties): They asked questions about four types of things:
    • Physical: Is it made of wood? Is it round?
    • Taxonomic: Is it a mammal? Is it a vehicle? (This requires knowing facts, not just seeing).
    • Functional: Can it be used to drive somewhere? Is it breakable?
    • Relational: Is it hanging on the wall? Is it next to a couple?
  • The Difficulty Levels (Reasoning Complexity):
    • Level 1 (Spotting): "How many dogs do you see?" (Just look and count).
    • Level 2 (Deducing): "How many things here are used for transportation?" (You have to look at a bike, a car, and a plane, realize they are all vehicles, and then count).
    • Level 3 (Imagining): "If we removed half the clocks, how many round things would be left?" (You have to imagine changing the picture and then do the math).
  • The Ingredients (Image Types): They used three kinds of pictures:
    • Photographs: Real, messy, real-world photos.
    • Animated: Clean, cartoon-style drawings.
    • AI-Generated: Pictures made by other AI, which sometimes look weird or impossible (like a glass floating in mid-air).

3. The Results: The Robots Are Still Learning

The researchers tested 12 of the smartest robots available. Here is what happened:

  • The Score: Even the best robot only got about 40% of the counting questions right and 70% of the comparison questions right. Humans, by contrast, got about 74% right.
  • The "Under-Counting" Habit: The robots had a bad habit of guessing "zero" or small numbers. If there were 8 items, they often guessed 2 or 3. It's like a student who is afraid to guess a high number, so they always pick the lowest one.
  • The "Real World" Struggle: The robots did great on cartoons and AI-generated images. But when they looked at real photographs, their scores dropped significantly. The messy lighting and hidden objects confused them.
  • The "Magic" Failure: The robots struggled the most with the "What if?" questions (Counterfactuals). If you asked, "What if we removed the lamp?", the robot often got confused about what was left. It's like asking a child to imagine a world without gravity; they get stuck on the current reality.

4. The Big Difference: Hallucinations vs. Confusion

The researchers looked closely at why the robots and humans got answers wrong. They found a funny difference:

  • Humans mostly got stuck on ambiguity. For example, "Is a half-wood, half-metal chair 'wood' or 'metal'?" Humans argue about the definition.
  • Robots mostly made "Unnatural" mistakes. They would count a glass as a metal object, or count a chair that wasn't even there. It's like a robot seeing a shadow and thinking it's a real dog. They "hallucinate" things that don't exist.

5. The New Kids on the Block

The researchers also tested some brand-new, super-smart "reasoning" models (like GPT-o3). These new models did much better, scoring around 52%. They made fewer "hallucination" mistakes. However, they still couldn't beat humans, and they still struggled with the messy real-world photos.

The Bottom Line

This paper tells us that while our AI robots are getting better at "seeing" and "thinking," they are still very different from humans.

  • Humans are great at imagining "what if" scenarios and handling messy photos, but we get confused by vague definitions.
  • Robots are great at spotting patterns in clean cartoons, but they get lost in the real world, often inventing things that aren't there and failing to do simple "what if" math.

The researchers say we need to keep building better tests to help these robots learn to see the world as clearly as we do.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →