← Latest papers
🧬 biology

The Illusion-Illusion: Vision Language Models See Illusions Where There Are None

This paper reveals that current vision-language models suffer from "illusion-illusions," erroneously perceiving non-illusory images (such as genuinely crooked lines or differently sized circles) as optical illusions, thereby exposing fundamental processing errors in their visual reasoning.

Original authors: Tomer Ullman

Published 2026-07-21
📖 1 min read☕ Coffee break read

Original authors: Tomer Ullman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Technical Summary: The Illusion-Illusion: Vision Language Models See Illusions Where There Are None

Problem Statement
Perceptual illusions serve as diagnostic tools in cognitive science by revealing gaps between reality and appearance, thereby illuminating underlying mental processing mechanisms. While previous research has investigated whether artificial systems fall prey to the same classical illusions as humans, this paper argues that testing classic illusions is insufficient for diagnosing current Vision Language Models (VLMs). The author posits that a more revealing diagnostic tool is the "illusion-illusion" (or "illusion-illusion"): images that possess the visual characteristics of an illusion but lack the actual perceptual trickery. In these cases, the visual reality matches the appearance (e.g., lines that are genuinely different lengths, a duck that is genuinely a duck), yet they are neighbors to famous illusions. The core problem addressed is whether current VLMs mistakenly perceive these "illusion-illusions" as illusions, thereby revealing a failure to distinguish between genuine perceptual anomalies and straightforward visual facts.

Methodology
The study employed a comparative evaluation framework involving ten representative visual illusions covering diverse perceptual phenomena (e.g., Müller-Lyer arrows, Ebbinghaus circles, Kanizsa triangles, checker shadows). For each classic illusion, the author constructed three types of stimuli:

  1. The Illusion: The original image designed to elicit a perceptual error.
  2. The Illusion-Illusion: A modified version where the "trick" is removed, and the visual properties are literal (e.g., arrows with genuinely different shaft lengths, a clear image of a duck).
  3. The Control: Simplified versions designed to verify basic competency (e.g., identifying a single triangle).

The evaluation tested eight current VLMs: GPT-4o, Claude 3, Gemini Pro Vision, miniGPT, Qwen-VL, InstructBLIP, BLIP2, and LLaVA-1.5. The models were presented with images and specific prompts targeting the illusion's subject (e.g., "Which is longer, the blue line or the red line?"). A secondary condition involved pre-pending the phrase "In the following visual illusion" to the prompts to test the influence of context.

Performance was scored as binary (0 or 1) based on alignment with human-like responses. For classic illusions, a score of 1 indicated reporting the illusory percept or acknowledging the illusion. For illusion-illusions and controls, a score of 1 indicated reporting the literal truth as a human would. The author notes this scoring is lenient, erring on the side of over-stating model performance.

Key Results
The results indicate a significant divergence between human perception and current VLM capabilities:

  • Failure on Illusion-Illusions: Leading models (GPT-4o, Claude 3, Gemini Pro), which successfully recognize or report classic illusions, also frequently misclassify illusion-illusions as illusions. For instance, when presented with lines that are genuinely different lengths, these models often claim they are the same length (mimicking the illusion) or describe them as an illusion.
  • Contextual Fragility: When the prompt explicitly labeled an image as a "visual illusion," the performance of the top three models degraded significantly on illusion-illusions and controls. They overwhelmingly reported these non-illusory images as illusions, suggesting their responses are heavily biased by the textual prompt rather than visual evidence.
  • Control Failures: Contrary to the hypothesis that controls would serve as a baseline for competency, models like GPT-4o, Gemini Pro, and Claude 3 frequently reported simple control images (e.g., a single triangle) as illusions. This suggests that for these models, the concept of "illusion" is being applied broadly to images that merely resemble the category of illusion, rather than being a specific perceptual judgment.
  • Model Behavior: The behavior of the top models does not align with a hypothetical "Model B" (which would recognize illusions but not illusion-illusions). Instead, they exhibit a pattern closer to "flat" performance or a tendency to hallucinate illusions where none exist, particularly when primed by the word "illusion."

Key Contributions
The paper introduces the concept of "illusion-illusions" as a novel diagnostic tool for artificial systems. By inverting the standard use of illusions, the author demonstrates that current VLMs suffer from a specific failure mode: they apply the pattern of illusion recognition to images that do not warrant it. This work highlights that the failure to distinguish between an illusion and an illusion-illusion is a more profound indicator of processing errors than the mere ability to detect a classic illusion. The study provides a dataset and evaluation protocol showing that models often fail to ground their responses in the literal visual reality when the context suggests an illusion is present.

Significance and Claims
The author claims that these failures are not merely quirks but indicative of broader processing errors in current VLMs. The paper suggests that when a model fails on an illusion-illusion, it implies the system is not engaging in thoughtful recognition of "ordinary" input. Instead, the model likely relies on low-level similarity matching to training data, associating the visual features of an image with the concept of an illusion found in its training corpus, rather than mimicking the specific perceptual failures of human biology.

The significance of this work lies in its challenge to the assumption that a model's ability to identify an illusion proves it possesses human-like perceptual understanding. The author argues that falling for an illusion is a human trait, but falling for an "illusion-illusion" (where the trick is absent) is a sign of a system that is "tricked" by the idea of an illusion rather than the visual data itself. The paper concludes that these failures suggest current systems are not yet robust in distinguishing between the appearance of a problem and the reality of the input, a limitation that extends beyond vision to other modalities and domains.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →