← Latest papers
🤖 AI

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

This paper introduces See2Think, a unified evaluation framework comprising a 1,200-sample benchmark and a Visual Action-of-Thought protocol, to reveal that while multimodal models exhibit behavioral dependence on intermediate visual states, their reasoning accuracy is heavily constrained by rendering fidelity and varies significantly across models and environments.

Original authors: Siyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou, Jingyu Chen, Chenhao Ji, Shuo Cao, Yongheng Zhang, Haoze Liu, Siyu Zhang, Xiwen Gu, Yihao Liu, Alex Jinpeng Wang

Published 2026-07-30
📖 4 min read☕ Coffee break read

Original authors: Siyu Yan, Zhuoran Yan, Haiying Xu, Panhao Zhou, Jingyu Chen, Chenhao Ji, Shuo Cao, Yongheng Zhang, Haoze Liu, Siyu Zhang, Xiwen Gu, Yihao Liu, Alex Jinpeng Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to solve a tricky puzzle, but instead of just looking at the picture, you are allowed to grab a pencil, draw lines, circle clues, and sketch out your thinking on a piece of paper. This is the world of Multimodal Large Language Models—super-smart computer programs that can read text and "see" images. For a while, these models have been getting better at "thinking out loud" by writing down their steps, a trick called Chain-of-Thought. But many puzzles, especially those involving 3D spaces or complex geometry, are hard to solve with words alone. Humans naturally draw diagrams to help us think, so scientists wondered: Can these AI models actually use intermediate drawings and sketches to solve problems, or are they just pretending to?

This is the big question researchers are asking. If an AI says, "I need to draw a line here to figure this out," does it actually use that new drawing to find the answer? Or does it just ignore the drawing and guess the answer anyway? It's like asking if a student is actually using their scratch paper to do the math, or if they are just scribbling nonsense and hoping the final answer comes from magic. Understanding this is crucial because if the models aren't truly using the visual tools they create, we might be wasting time building features that don't actually make them smarter.

Enter See2Think, a new study that acts like a detective for AI reasoning. The researchers built a massive testing ground called See2ThinkBench, containing 1,200 different puzzles ranging from 2D geometry and chemistry diagrams to 3D robot scenes and real-world physics problems. They didn't just ask the models to give an answer; they set up a special protocol called Visual Action-of-Thought (VAoT). Think of this as a game where the AI has to take turns: it thinks, it decides to "draw" something (like highlighting a specific angle or cropping a part of the image), a separate tool actually draws it, and then the AI has to look at that new drawing to continue its reasoning.

The team tested four different powerful AI models to see how they handled this game. They found that there is no "one size fits all" winner. Sometimes, just thinking in words (without drawing) worked best; other times, actually seeing the drawing helped. But here is the twist: the models were surprisingly good at deciding what to draw (picking the right clues), but they often failed at using the drawing correctly. In fact, the study found that the biggest bottleneck wasn't choosing the right action, but rather the "faithfulness" of the drawing itself—did the tool actually draw what the AI asked for?

Perhaps the most fascinating discovery came from a "trick" the researchers played. They intentionally gave the models "corrupted" drawings—pictures that looked like the AI asked for but were actually wrong or misleading. When this happened, the models' performance dropped significantly, especially in 3D scenes, where accuracy fell by over 10 percentage points. This proves that the models were genuinely relying on the visual state they saw, even if that state was broken. However, the study also suggests that just because a model depends on the drawing doesn't mean the drawing actually helped it get the right answer. Sometimes, the models were so dependent on the visual feedback that even a bad drawing confused them enough to fail, showing that "using" a visual state and "benefiting" from it are two very different things.

In short, the paper reveals that while AI models are getting better at the idea of visual reasoning, they still struggle with the messy reality of executing and trusting their own sketches. They are like students who know exactly which part of the diagram to look at, but often get lost when they try to actually read the new sketch they made. The researchers conclude that we need to stop just looking at the final answer and start diagnosing the whole process—checking if the drawing was relevant, if it was drawn correctly, and if the model actually used it to think. Until we fix these bottlenecks, the "thinking with images" might still be more of a performance than a genuine superpower.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →