LookBack: Where and How to Score LVLM Responses via Visual Reference Usage
This paper introduces LookBack, a training-free method that improves LVLM response scoring by augmenting token likelihood with a visual lookback score to better detect hallucinations and select grounded outputs, addressing the limitations of existing confidence-based metrics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are talking to a super-smart robot that has both eyes and a brain. This robot, called a Large Vision-Language Model (or LVLM for short), can look at a picture and then write a story or answer a question about it. It's like having a friend who can see the world and describe it perfectly. But here's the catch: sometimes this robot gets too confident in its own voice. It might describe a scene with such perfect grammar and flow that it sounds totally real, even if it's making things up that aren't in the picture at all. This is called "hallucinating."
To fix this, scientists usually try to pick the best answer from a list of many guesses the robot makes. They often use a simple trick: they pick the answer that sounds the most confident or likely based on the robot's own internal math. Think of it like a teacher picking the essay that sounds the most fluent. But what if the robot is just a smooth talker who is confidently wrong? That's the problem this paper tackles: how do we tell if the robot is actually looking at the picture, or if it's just daydreaming while sounding smart?
The researchers behind this study, from Yonsei University, realized that the robot's usual "confidence meter" is broken when it comes to pictures. They found that even if you take the picture away and just ask the robot to guess, it still gives the same high confidence scores to its answers. It's like a student who can recite a history essay perfectly from memory, but if you ask them about a specific photo in their textbook, they might confidently describe a dinosaur in a classroom because they are just good at guessing words, not looking at evidence.
To solve this, the team invented a new method called LOOKBACK. Imagine you are grading that student's essay again. Instead of just checking if the sentences flow well, you ask: "Did you actually look at the photo when you wrote this word?" LOOKBACK is a special tool that checks every single word the robot writes to see if the robot's "eyes" (its internal attention) were actually looking at the picture at that exact moment.
Here is how it works in a fun way:
- The Confidence Check: First, LOOKBACK checks how sure the robot is about a word, just like the old methods did.
- The Lookback Check: Then, it asks, "Did you look at the picture to say this?" If the robot says "panda" and its internal eyes were staring right at a panda in the image, that word gets a huge boost. If the robot says "panda" but was just guessing because it loves the word "panda," it gets a penalty.
- The Final Score: The tool combines these two checks. It gives the highest score to answers that are not only fluent and confident but also prove they were actually looking at the visual evidence.
The team tested this idea on four different challenges involving pictures and questions, using three different smart robot brains. They found that LOOKBACK was much better at picking the correct answer than the old methods. In fact, when they had the robots generate 25 different guesses and used LOOKBACK to pick the winner, it consistently chose the right answer more often than any other method they tried.
What's really cool is that LOOKBACK doesn't need to learn anything new or use any extra robots to help. It just uses the signals the main robot is already sending while it thinks. It's like giving the robot a mirror so it can check its own work without needing a teacher to stand over it. The study suggests that by simply asking the robot to "look back" at the picture while it speaks, we can stop it from confidently lying about what it sees, making these AI friends much more reliable for real-world tasks.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.