ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination
ReGround is a two-stage framework that restores visual grounding in Vision-Language Models during multi-step reasoning by teaching them to autonomously self-diagnose grounding failures and selectively re-examine visual evidence, achieving significant performance gains without architectural modifications or external tools.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a tricky puzzle, but instead of just looking at the picture on the box, you have a super-smart robot friend who can talk to you while they look. This is what happens with "Vision-Language Models" (VLMs). These are AI systems that can see an image and understand a question about it, then chat through the steps to find the answer. Think of them as a detective who can look at a crime scene photo and a witness statement simultaneously.
However, there's a catch. When the detective has to take many steps to solve a complex case—like a long math problem or a science puzzle—they sometimes start to forget the photo. As they talk more and more, their brain gets filled with words, and the image in their mind starts to fade. They might start guessing based on what they think usually happens, rather than what is actually in the picture. This is called "visual grounding decay." It's like a detective who stops looking at the evidence and just starts guessing the culprit's name because it sounds right. Scientists care about this because if the AI forgets the picture, it can make silly mistakes, like saying a blue car is red just because it's talking too much about colors.
Enter ReGround, a new method that acts like a "reality check" for these AI detectives. The researchers found that when an AI gets stuck in a long chain of reasoning, it often loses its grip on the image. But simply telling the AI to "look again" isn't enough; in fact, if you just say "look again" without a specific reason, the AI might get confused and make more mistakes. It's like a student who, when told to "check your work," just panics and changes a correct answer to a wrong one because they don't know what to look for.
The paper shows that the secret sauce is Self-Diagnosis. ReGround teaches the AI to first stop and ask itself, "Wait, did I actually look at the picture for this step, or am I just guessing?" If the AI realizes it's drifting away from the image, it triggers a special process. It then re-injects the original image into its memory, but this time, it comes with a specific note from the AI itself, like, "Hey, look at the angle of this triangle again," or "Check if that number matches the chart."
The researchers tested this on eight different benchmarks, which are like standardized tests for AI. They found that when the AI used this "diagnose and re-check" method, it got significantly better at solving hard visual problems. For example, on a test called HallusionBench (which is designed to trick AI into hallucinating things that aren't there), the AI's score jumped by 6.8 points. This is a big deal because it means the AI is less likely to make up facts.
Crucially, the paper rules out a few ideas. It proves that just having the AI talk to itself (text-only reflection) isn't enough; the image must be shown again. It also shows that a generic "look again" command is actually harmful. The AI needs a specific, targeted reason to look back. The study suggests that the quality of this "self-diagnosis" is the most important part. If the AI can learn to diagnose its own confusion well, it can recover most of the benefits of having a super-smart teacher help it, without needing any external tools or changing the AI's basic brain structure.
In short, ReGround teaches AI to be a better detective: not just by looking at the evidence again, but by knowing exactly what to look for and when to stop guessing and start checking. It's a way to keep the AI's eyes on the prize, even when the reasoning gets long and complicated.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.