← Latest papers
🤖 AI

Thinking with Visual Grounding

This paper introduces "visually grounded thinking," a framework where vision-language models interleave natural language reasoning with explicit visual groundings, which, when trained via a scalable synthesis pipeline and grounding-aware reinforcement learning, significantly enhances performance on counting and spatial reasoning tasks, allowing smaller models to rival much larger counterparts.

Original authors: Junkai Zhang, Yihe Deng, Kai-Wei Chang, Wei Wang

Published 2026-06-16
📖 4 min read☕ Coffee break read

Original authors: Junkai Zhang, Yihe Deng, Kai-Wei Chang, Wei Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are taking a math test, but instead of numbers, the questions are about pictures. You have a smart assistant (an AI) that tries to solve these picture puzzles by talking to itself out loud.

The Problem: The "Ghost" Reasoning
Currently, these AI assistants are good at talking. They can say, "I see a red car near the entrance, so the answer is A." But here's the catch: they are just saying they see it. They aren't actually pointing to the red car in the picture. It's like a student taking a test who writes down the right answer but can't show their work. If you ask, "How do you know that's the red car?" the AI might just say, "I just know," or it might be guessing. Because the AI isn't forced to point to the evidence, it's hard to tell if it's actually looking at the picture or just making things up.

The Solution: "Visual Grounding"
The authors of this paper, Junkai Zhang and his team, came up with a new way for the AI to think called "Visually Grounded Thinking."

Think of it like this:

  • Old Way: The AI says, "There is a black laptop on the table." (It's just words).
  • New Way: The AI says, "There is a black laptop on the table."

Every time the AI mentions an object (like a "brown chair" or a "blue cup"), it has to drop a digital pin (a point) or draw a box around that specific object in the image. It's like the AI is holding a laser pointer and saying, "I am thinking about this specific thing right here."

How They Taught the AI to Do This
Teaching an AI to point accurately is hard because you can't just ask it to "be precise." So, the researchers built a special factory (a data pipeline) to create training examples:

  1. The Detective: They used a very smart AI to solve picture puzzles and write down the steps.
  2. The Highlighter: They used a tool called SAM3 (think of it as a super-precise digital highlighter) to automatically find the exact shapes of the objects the detective mentioned.
  3. The Teacher: They took those exact shapes and turned them into "points" or "boxes" to create a perfect answer key.
  4. The Reward System: They trained the AI using a game-like system. If the AI got the answer right and pointed to the right spot, it got a high score. If it got the answer right but pointed to the wrong spot (or didn't point at all), it got a lower score. This forced the AI to learn that "thinking" and "pointing" must happen together.

What Happened When They Tested It?
They tested this new "pointing" AI on two types of tasks:

  1. Counting: "How many donuts are in this picture?"
    • Result: The AI got much better at this. Pointing to each donut helped it count accurately without getting confused by similar-looking items.
  2. Spatial Reasoning: "Which object is farthest to the left of the man?"
    • Result: This was the big surprise. A small AI model (4 billion "brain cells") that learned to point became just as good, and sometimes even better, than a massive AI model (27 billion "brain cells") that didn't learn to point.

The Takeaway
The paper shows that when AI models are forced to tie their thoughts to the actual image—like a student who must show their work by circling the numbers they used—they become much smarter.

  • For counting: Just pointing a dot at an object is enough.
  • For complex spatial puzzles: Drawing a box around the object helps the AI understand size and shape better, but even just pointing works surprisingly well.

In short, the paper proves that thinking is better when you can point to what you're thinking about. It turns "magic guessing" into "verifiable evidence."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →