← Latest papers
💻 computer science

Comparative Evaluation of Machine Translation Systems on Images with Text

This paper presents a comparative evaluation of machine translation systems for images containing text, demonstrating that multi-modal large language models outperform both modular OCR-based pipelines and end-to-end approaches in terms of translation quality and contextual understanding.

Original authors: Blai Puchol, Sergio Gómez González, Miguel Domingo, Francisco Casacuberta

Published 2026-05-29
📖 4 min read☕ Coffee break read

Original authors: Blai Puchol, Sergio Gómez González, Miguel Domingo, Francisco Casacuberta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a photo of a street sign in a foreign country. You want to know what it says, but you don't speak the language. This paper is like a race between three different teams of "digital translators" to see which one can read that sign and tell you its meaning most accurately.

Here is how the paper breaks down the competition, using simple analogies:

The Three Contenders

The researchers tested three different ways to solve the problem:

  1. The Assembly Line (Modular Pipeline):
    Think of this as a factory with two distinct workers.

    • Worker A (The Eyes): First, a specialized robot looks at the picture and writes down exactly what the letters say. This is called OCR (Optical Character Recognition).
    • Worker B (The Translator): Then, a second robot takes that written list of letters and translates it into your language.
    • The Paper's Finding: This team used the best "eyes" (a tool called docTR) and the smartest "translators" (AI models like Llama and EuroLLM).
  2. The Super-Brain (Multi-Modal Large Language Models):
    This is like hiring a genius who can see the picture and read the text at the same time. Instead of passing the work from one person to another, this single AI looks at the image, understands the context, and instantly tells you the translation.

    • The Paper's Finding: The "Super-Brains" (specifically Google's Gemini models) were the clear winners. They didn't just translate words; they understood the whole picture better than anyone else.
  3. The Direct Painter (End-to-End Model):
    This is a robot that tries to take the original photo and magically paint the new words directly onto the image, replacing the old ones. It skips the "reading" and "translating" steps and tries to do everything in one giant leap.

    • The Paper's Finding: This approach struggled the most. It was like trying to paint a perfect portrait without ever sketching the outline first. It made more mistakes than the other two teams.

The Race Track (The Experiment)

To make sure the race was fair, the researchers didn't use messy, real-world photos with blurry lights or crooked signs. Instead, they created a "practice track" using computer-generated images. They took sentences in German, French, and Romanian, typed them out in black font, and placed them on colorful backgrounds with different rotations.

This ensured that every AI team faced the exact same conditions, making it easy to see who was actually the best translator and who was just guessing.

The Results: Who Won?

When the dust settled, the results were clear:

  • The Winners: The Multi-Modal Models (The Super-Brains) took first place. They were the most flexible and understood the context of the text best. They didn't get confused by the layout of the image; they just "got it."
  • The Runner-Up: The Assembly Line (Modular Pipeline) came in second. It did a great job, especially when they used the best "eyes" and the smartest "translators" available. It proved that breaking a big problem into smaller, specialized steps works very well.
  • The Loser: The Direct Painter (End-to-End) came in last. It struggled to get the translation right, suggesting that trying to translate an image directly without first reading the text is still a very hard challenge for computers.

The Catch (Limitations)

The paper also admits a few things about the race:

  • The Track Was Fake: The images were perfectly generated by a computer. In the real world, photos are messy (blurry, dark, or covered in rain). The winners might do even better or worse when facing real-life chaos.
  • The Languages Were Limited: The race only included a few European languages. We don't know if these same winners would be just as good with languages that use different scripts (like Chinese or Arabic) or very rare languages.
  • The Scorecard: The judges used computer formulas to grade the answers. While these formulas are standard, they don't measure how "human" or natural the translation feels.

The Bottom Line

The main takeaway is simple: If you want to translate text in an image today, the best strategy is to either use a "Super-Brain" AI that sees and reads simultaneously, or to use a high-quality "Assembly Line" that separates reading from translating. Trying to do it all in one giant leap (the Direct Painter) isn't quite ready for prime time yet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →