← Latest papers
🤖 AI

HueManity: Probing Fine-Grained Visual Perception in MLLMs

The paper introduces HueManity, a large-scale benchmark using Ishihara-style images to reveal that state-of-the-art Multimodal Large Language Models (MLLMs) exhibit a critical deficit in fine-grained visual perception compared to humans and specialized models, despite their strong high-level reasoning capabilities.

Original authors: Rynaa Grover, Jayant Sravan Tamarapalli, Sahiti Yerramilli, Nilay Pande

Published 2026-02-03
📖 4 min read☕ Coffee break read

Original authors: Rynaa Grover, Jayant Sravan Tamarapalli, Sahiti Yerramilli, Nilay Pande

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart robot that can read books, write poems, and describe complex scenes in a painting. You might think, "If it's so smart, it can surely see everything clearly, right?"

The paper "HueManity" says: "Not so fast."

The researchers built a special test to see if these "Multimodal Large Language Models" (MLLMs)—the AI brains behind many chatbots—can actually see the tiny details, or if they are just guessing based on what they've read before.

Here is the breakdown of their findings using simple analogies:

1. The Test: The "Invisible Ink" Puzzle

The researchers created a massive library of 83,850 images. Each image looks like a colorful, messy pile of dots, similar to the Ishihara color blindness tests you might have seen in a doctor's office.

  • The Trick: Hidden inside the messy dots are two letters or numbers (like "42" or "X7").
  • The Challenge: To find them, you have to ignore the noise and spot the subtle color differences that form the shape of the letters.
  • The Goal: It's not about "thinking" or "reasoning." It's purely about seeing. Can the AI pick out the needle in the haystack?

2. The Results: Humans vs. AI

The researchers ran this test on humans, a standard computer vision program (ResNet50), and nine of the smartest AI models available today (like GPT-4, Claude, and LLaVA).

  • Humans: Were almost perfect. They saw the numbers and letters instantly (99% and 93% accuracy).
  • Standard Computer Vision: Also did great (96% and 94%). It's like a trained eye that knows exactly what to look for.
  • The "Smart" AI Models: They failed miserably.
    • The best AI model only got 33% of the simple number tests right.
    • On the harder letter/number tests, the best AI got only 3% right.
    • Most other models got 0% to 1%.

The Analogy: Imagine asking a genius who has read every book in the library to find a specific word written in invisible ink on a page. The genius might know the concept of the word, but they can't actually see the ink on the page. They are "hallucinating" answers instead of seeing the truth.

3. Why Did the AI Fail?

The paper suggests the AI isn't "dumb," but it has a specific blind spot:

  • Too Much Focus on "Big Picture": These AIs are trained to understand the meaning of an image (e.g., "This is a cat sitting on a mat"). They are so good at the big picture that they ignore the tiny, messy details (the specific colors and shapes of the dots).
  • Relying on Guessing: When the AI can't see the dots, it tries to guess based on language patterns. It's like a student who doesn't know the math problem but guesses an answer because it sounds like something they've heard before.
  • The "Squinting" Effect: Interestingly, when the researchers made the images blurrier (lower resolution), one AI model actually got better. It seems the AI needs the "noise" to be smoothed out to guess the shape, rather than actually seeing the details.

4. The "Magic Trick" That Didn't Work

The researchers tried to "teach" the AI how to do this test:

  • Showing Examples: They showed the AI a few examples of the puzzle first. It didn't help.
  • Fine-Tuning: They tried to retrain the AI specifically on these puzzles. It didn't help. In fact, it made the AI forget how to read simple text!

The Conclusion: The problem isn't that the AI hasn't been taught enough; the problem is likely how the AI is built. It's like trying to teach a fish to climb a tree by giving it more books on climbing. The fish (the AI) just isn't built for that kind of vision.

Why Does This Matter?

The paper argues that we can't trust these AIs in situations where seeing clearly is critical.

  • If an AI is driving a car, it needs to see a faint crack in the windshield or a small sign in the rain, not just "guess" that there is a road.
  • If an AI is looking at medical scans, it needs to spot a tiny anomaly, not just describe the general shape of the organ.

In short: These AI models are excellent storytellers and great at understanding concepts, but they are currently terrible at fine-grained vision. They are like a person who can describe a painting perfectly but can't find the hidden signature in the corner.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →