← Latest papers
💻 computer science

Hidden Meanings in Plain Sight: RebusBench for Evaluating Cognitive Visual Reasoning

The paper introduces RebusBench, a benchmark of 1,164 rebus puzzles designed to evaluate the cognitive visual reasoning of Large Vision-Language Models, revealing that despite their strong perceptual capabilities, current state-of-the-art models severely struggle with the multi-step neurosymbolic reasoning required to solve these puzzles, achieving less than 10% exact match accuracy regardless of model scale or in-context learning.

Original authors: Seyed Amir Kasaei, Arash Marioriyad, Mahbod Khaleti, MohammadAmin Fazli, Mahdieh Soleymani Baghshah, Mohammad Hossein Rohban

Published 2026-04-03
📖 4 min read☕ Coffee break read

Original authors: Seyed Amir Kasaei, Arash Marioriyad, Mahbod Khaleti, MohammadAmin Fazli, Mahdieh Soleymani Baghshah, Mohammad Hossein Rohban

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart robot that can look at a picture and describe exactly what it sees. If you show it a photo of a cat sitting on a mat, it will happily tell you, "There is a cat on a mat." This robot is great at seeing.

But what happens if you show the robot a picture that isn't just a photo, but a visual riddle?

The Problem: The Robot Can See, But Can't "Get" the Joke

The paper introduces a new test called RebusBench. A "rebus" is a puzzle where pictures and letters are mixed together to hide a common phrase.

Think of it like this:

  • The Picture: You see a red letter "E" and two words saying "GO."
  • The Robot's View: "I see a red 'E' and two 'GO's." (This is what the robot actually says).
  • The Human View: "Wait, 'Red E' sounds like 'Ready', and 'Two Gos' sounds like 'To Go'. The answer is 'Ready to go'!"

To solve this, you can't just look at the pixels. You have to:

  1. Notice the color and the repetition.
  2. Remember that "Red E" sounds like "Ready."
  3. Connect those sounds to a common idiom.

This is called System 2 thinking—slow, deliberate, creative reasoning. It's the difference between recognizing a face (System 1) and solving a mystery (System 2).

The Experiment: The "Giant Brain" Fails the Riddle Test

The researchers took the smartest, biggest AI models available (like Qwen, InternVL, and LLaVA) and gave them 1,164 of these riddles. They tried everything to help the models:

  • Making the models bigger: They used models with billions of parameters (like upgrading from a bicycle to a rocket ship).
  • Giving them examples: They showed the models a few solved riddles first (like giving a student a practice test).

The Result? The models failed miserably.

  • Even the biggest models only got about 5% to 8% of the answers right.
  • Making the models bigger didn't help much.
  • Giving them examples didn't help much either.

The Analogy: The Library vs. The Librarian

Imagine these AI models are like massive libraries containing every book ever written. They have all the words, all the pictures, and all the facts.

  • What they are good at: If you ask, "What is in this picture?" they can pull out the book that says "Cat" and show you the picture of a cat.
  • What they are bad at: If you ask a riddle, they get stuck. They have all the ingredients (the word "Red," the word "E," the word "Go"), but they lack the chef who knows how to mix them together to make a new dish.

They have the ingredients (vision and language), but they are missing the cooking recipe (the cognitive "glue" that connects the dots).

Why Does This Matter?

The paper argues that we are hitting a wall. We keep making AI bigger and giving it more data, but it still can't "think" in the way humans do when solving puzzles.

  • Current AI: "I see a red E and two Gs." (Literal)
  • Human Intelligence: "Red E + Two Gs = Ready to Go!" (Abstract)

The researchers call this a "cognitive gap." It suggests that simply adding more computing power isn't enough to teach AI how to be creative or solve riddles. We need a new kind of "glue" to help these models connect what they see with what they know in a clever way.

The Bottom Line

The paper is a wake-up call. It says: "Our AI is amazing at describing the world, but it's terrible at understanding the hidden jokes and riddles of the world."

Until we figure out how to give AI that "aha!" moment, it will remain a very smart parrot that can repeat facts but can't quite solve the puzzle.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →