← Latest papers
💬 NLP

Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions

This paper demonstrates that Vision-Language Models (VLMs) performing OCR on Ancient Greek texts often rely on language priors to generate fluent but visually ungrounded errors, a failure mode that persists even with decode-time interventions and differs significantly from the noise patterns of traditional OCR systems.

Original authors: Antonia Karamolegkou, Nicolas Angleraud, Benoît Sagot, Thibault Clérice

Published 2026-05-28
📖 5 min read🧠 Deep dive

Original authors: Antonia Karamolegkou, Nicolas Angleraud, Benoît Sagot, Thibault Clérice

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to read a very old, handwritten letter written in Ancient Greek. The ink is faded, the letters are strange, and the page is covered in messy notes in the margins. You have two helpers to read it for you:

  1. The "Old-School Scanner" (Traditional OCR): This is like a strict, literal robot. It looks at the ink on the page and says, "I see a squiggle that looks like an 'A', so I will write an 'A'." If the ink is too blurry, it might write a random symbol or a typo. It doesn't guess; it just reports what it sees, even if the result looks messy.
  2. The "Smart AI Assistant" (Vision-Language Model): This is like a brilliant student who knows Ancient Greek perfectly but is a bit lazy about looking closely at the paper. When they see a blurry spot, they think, "Hmm, based on the sentence so far, the next word should be 'king'." So, they write "king," even if the ink on the page actually looked more like "dog."

This paper, titled "Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions," investigates what happens when these two helpers try to read difficult, low-resource Ancient Greek texts.

Here is what the researchers found, using simple analogies:

1. The "Fluent Lie" vs. The "Messy Truth"

When the Old-School Scanner makes a mistake, it usually looks like a typo or a jumble of letters (e.g., "kng"). It's obvious something went wrong.

However, when the Smart AI Assistant makes a mistake, it often writes a perfectly real, fluent Ancient Greek word that just isn't what was written on the page.

  • The Analogy: Imagine the AI is reading a recipe. If the ingredient list is smudged and says "flour," but the AI knows this recipe usually calls for "sugar," it might confidently write "sugar." To a quick glance, the sentence looks perfect. But it's actually a lie. The AI is relying on its memory of how recipes usually go (its "language prior") rather than looking at the specific paper in front of it.

2. The "Magic Ink" Test

To prove this, the researchers played a trick. They took the text, scrambled the letters inside the words (turning real words into nonsense gibberish), and then showed this gibberish to the helpers.

  • The Old-School Scanner: It looked at the nonsense and wrote down the nonsense. It stayed faithful to the visual evidence, even though the result was garbage.
  • The Smart AI Assistant: It looked at the nonsense, got confused, and then "fixed" it. It rewrote the gibberish back into real, fluent Greek words. It ignored the visual evidence (the scrambled letters) and relied entirely on its internal knowledge of the language.

3. Not All "Smart" AIs Are the Same

The researchers found a surprising twist. Not all AI assistants behave the same way.

  • The Generalist AI: These are the big, famous models used for many tasks. Even when they were wrong, they were still "looking" at the picture. They were trying to read the ink, but their brain was just too eager to guess the right word.
  • The Specialist AI: There was one model built specifically for reading documents (called OlmOCR). This one was the worst offender. When it made a mistake, it was almost as if it didn't look at the picture at all. It was purely guessing based on what it thought the text should say. It was like a student who closes their eyes and recites the textbook from memory, ignoring the actual exam paper.

4. Can We Fix the AI?

The researchers tried several "patches" to force the AI to look at the picture instead of guessing:

  • The "Vocabulary Fence": They tried to tell the AI, "You can only write Greek letters." Result: This made things much worse. The AI got confused and started writing gibberish because it couldn't use its usual tricks.
  • The "Contrast Test": They tried to show the AI the picture and a blurry version of the picture at the same time to force it to focus on the details. Result: This helped a little bit for some models, but not for the specialist one.
  • The "Post-Game Edit": They let the AI write its answer first, and then had a second AI (a text-only editor) fix the mistakes. Result: This worked well to clean up the text, but it didn't fix the root problem. The first AI still ignored the picture; the second AI just fixed the lie after it was told.

The Big Takeaway

The main lesson is that just because an AI's answer sounds perfect and fluent, it doesn't mean it actually read the document.

In the world of ancient history and rare documents, this is dangerous. If a digital library uses these AI models to scan old Greek books, the AI might silently replace a rare, obscure word with a common, "plausible" one. A historian reading the digital version would never know the change happened because the sentence still looks grammatically perfect.

The paper concludes that we need new ways to test AI, not just by checking if the final score is high, but by checking if the AI is actually "grounded" in the visual evidence, or if it's just guessing based on what it thinks should be there.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →