← Latest papers
🤖 AI

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms

This paper investigates whether vision-language models exhibit human-like visual search behaviors by using reasoning token counts as a proxy for reaction time, finding that while models replicate key human signatures like flat feature search costs and rising conjunction costs, they also reveal distinct cognitive divergences in target-present slopes, enumeration accuracy, and adaptive deliberation strategies.

Original authors: Farahnaz Wick

Published 2026-06-25
📖 5 min read🧠 Deep dive

Original authors: Farahnaz Wick

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a game of "Where's Waldo?" in a crowded picture. If you are looking for a red hat in a sea of blue hats, your brain spots it instantly, no matter how many people are in the crowd. But if you are looking for a red hat that is also shaped like a star (and everyone else is wearing either blue hats or red circles), your brain has to scan the picture one person at a time. The more people there are, the longer it takes you to find the target.

This is the classic "Visual Search" game that psychologists have used for decades to understand how human attention works. A new paper asks a fascinating question: Do AI models play this game the same way humans do?

Since AI doesn't have a heartbeat or a stopwatch, the researcher couldn't measure how long the AI "thought" about an image. Instead, they used a clever trick: they counted the number of "thinking tokens" the AI generated before giving an answer. Think of these tokens like the AI's internal monologue or its "scratchpad" notes.

  • Few tokens = The AI glanced at the image and said, "Easy, I see it." (Fast reaction).
  • Many tokens = The AI wrote a long list of reasons, checked the image again, and debated the answer. (Slow, effortful reaction).

Here is what the study found, using simple analogies:

1. The "Pop-Out" vs. The "Needle in a Haystack"

Human Behavior: When the target is obvious (a red hat among blue ones), humans find it instantly. When it's tricky (a red star among red circles), humans take longer as the crowd gets bigger.
AI Behavior: The smartest AI models did exactly the same thing.

  • For the easy "pop-out" tasks, the AI used a tiny, constant amount of thinking tokens, regardless of how many items were in the picture.
  • For the tricky "needle in a haystack" tasks, the AI's token count climbed as the picture got more crowded.
    The Takeaway: The AI's internal effort curve looks just like the human reaction-time curve. It suggests that when these AIs are "thinking," they are actually doing a serial, item-by-item search, just like a human eye scanning a room.

2. The "Empty Room" Reversal

Human Behavior: If you are looking for a specific item and it's not there, you have to check every single person in the crowd to be sure. This takes a lot of time. If the item is there, you can stop as soon as you find it. So, humans work harder when the answer is "No."
AI Behavior: The AI flipped this script. It worked harder when the answer was "Yes" and found the item, and it worked less when the answer was "No."
The Takeaway: It seems the AI has a strategy of "Find one, then justify it." Once it spots the target, it spends a lot of tokens double-checking and explaining its finding. But if it doesn't see one immediately, it might just guess "No" quickly without doing a full, patient sweep of the whole image.

3. The "Counting" Superpower

Human Behavior: If you ask a human to count 5 or 6 items in a busy picture, they often lose track or make mistakes. Our brains get tired of holding the count in our heads.
AI Behavior: The AI models were perfect at counting. Even in crowded pictures with many items, they didn't lose their count.
The Takeaway: Instead of getting tired and making errors like humans, the AI just spent more energy. It didn't drop the ball; it just grinded through the task with extra computing power. It traded "effort" for "accuracy."

4. The "Tilted Bar" Puzzle

Human Behavior: Humans are great at spotting a tilted bar among straight vertical bars (it "pops out"). But we are terrible at spotting a straight bar among tilted ones (it's hard).
AI Behavior: The AI models also found one direction harder than the other, but they expressed this difficulty differently depending on the model:

  • Model A (The Grind): Spent a massive amount of thinking tokens on the hard task but still got the right answer.
  • Model B (The Shortcut): Spent almost no thinking tokens on the hard task and simply got the answer wrong (hallucinating a tilted bar that wasn't there).
    The Takeaway: The same visual difficulty can be solved by "working harder" or by "giving up and guessing," depending on the specific AI's personality.

The Big Picture

The paper concludes that modern AI models have developed a "fingerprint" of human visual search. They know when to scan quickly and when to scan slowly. However, they are not human. They don't get tired in the same way, they stop searching differently, and they handle "hard" tasks by either grinding through them with extra effort or failing spectacularly.

The researcher argues that by watching how much an AI thinks (its token count) rather than just what it answers, we can see the hidden mechanics of its "vision." It's like listening to the hum of an engine to understand how a car works, rather than just looking at where it ends up.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →