← Latest papers
💬 NLP

Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding

This paper reveals a significant gap between the internal representations and generated responses of Large Vision Language Models in Visual Document Understanding tasks, demonstrating that task-relevant information is often more accessible in intermediate layers and that fine-tuning these layers effectively narrows this gap while improving overall performance.

Original authors: Haruka Kawasaki, Ryota Tanaka, Kyosuke Nishida

Published 2026-04-07
📖 4 min read☕ Coffee break read

Original authors: Haruka Kawasaki, Ryota Tanaka, Kyosuke Nishida

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: "Knowing" vs. "Saying"

Imagine you are taking a very difficult test about a complex document (like a tax form or a medical chart). You are a super-smart AI assistant.

The researchers in this paper discovered something surprising: Just because you get the right answer on the test doesn't mean you actually understood the question perfectly.

In fact, they found that the AI often "knows" the answer deep inside its brain (in its internal data), but when it tries to speak out loud (generate a text response), it fumbles and gives a wrong or incomplete answer. It's like a student who knows the math formula perfectly but writes down the wrong number on the final exam because they got nervous or distracted.

The Problem: The "Gap"

The paper calls this the "Gap between Internal Representations and Responses."

  • Internal Representation: The AI's internal "thoughts" or data processing.
  • Response: The actual text the AI types out for you.

Usually, we judge AI by how good its answers are. But this paper says, "Wait a minute! Let's peek inside the AI's brain while it's thinking."

They used a tool called Linear Probing. Think of this as a X-Ray Machine for the AI's brain. Instead of waiting for the AI to speak, the researchers put a small sensor on different parts of the AI's brain to see if the answer is already there.

The Discovery: The "Middle Layer" Surprise

The AI is built like a multi-story building with many floors (layers).

  1. Bottom Floors: Where the AI first sees the picture and reads the text.
  2. Top Floor: Where the AI decides what to say next.

The researchers expected that the "Top Floor" would have the clearest, most accurate answer. They were wrong.

They found that the answer was actually most clear and easy to find on the Middle Floors. By the time the information reached the Top Floor, it had gotten messy, diluted, or confused.

The Analogy:
Imagine a game of "Telephone" played in a factory.

  • The Bottom Floor is the raw material (the document image).
  • The Middle Floor is where the expert workers assemble the product. The product is perfect here.
  • The Top Floor is the shipping department. By the time the product gets to shipping, it's been handled so many times, wrapped in too much tape, and passed through too many hands that it arrives at the customer (you) slightly damaged or wrong.

The Solution: Fixing the Middle

Since the "perfect answer" exists in the middle floors but gets ruined on the way to the top, the researchers tried a new training method.

Instead of training the whole AI from top to bottom (which is expensive and doesn't fix the specific problem), they decided to only train the Middle Floors.

The Result:

  • The AI got better at answering questions.
  • The "Gap" between what it knew and what it said got much smaller.
  • It was also faster and cheaper to train this way.

Summary of Key Takeaways

  1. Don't just trust the answer: An AI can have the right answer hidden inside its "brain" even if it types the wrong thing.
  2. The "Sweet Spot" is in the middle: The most accurate information for visual documents isn't at the very end of the AI's processing; it's in the middle layers.
  3. Targeted training works best: Instead of retraining the whole AI, just giving a "tune-up" to the middle layers makes the AI smarter, more honest, and more efficient.

Why Does This Matter?

This is a big deal for safety and reliability. If we only look at the final answer, we might think an AI is failing when it's actually "thinking" correctly but just failing to communicate. By understanding this gap, we can build better AI systems that are more trustworthy, especially when dealing with important documents like legal contracts or medical records.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →