← Latest papers
🤖 AI

MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing

The paper introduces MULTITEXTEDIT, a controlled benchmark spanning 12 languages and a novel language fidelity metric, to demonstrate that current text-in-image editing systems suffer from significant cross-lingual degradation, particularly in script accuracy for non-Latin languages, despite maintaining global visual structure.

Original authors: Liwei Cheng, Zirui Song, Shibo Feng, Lunjie Zhou, Yixuan Guan, Dayan Guan

Published 2026-05-12
📖 4 min read☕ Coffee break read

Original authors: Liwei Cheng, Zirui Song, Shibo Feng, Lunjie Zhou, Yixuan Guan, Dayan Guan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a digital photo of a poster with the words "SALE" written on it. You want to use an AI tool to change those words to "SOLD" without messing up the background, the colors, or the picture of the shirt behind the text. This is called "text-in-image editing."

While these AI tools are getting very good at doing this in English, the paper MULTITEXTEDIT asks a simple but crucial question: What happens when we ask the AI to do the same thing in other languages?

Here is the breakdown of their findings, explained with everyday analogies:

1. The Problem: The "English-Only" Blind Spot

Think of current AI image editors like a chef who is a master at cooking Italian food but has never tried cooking Thai, Arabic, or Japanese cuisine. If you ask them to swap out an ingredient in a Thai dish, they might keep the plate and the table setting perfect, but they might serve you the wrong spice or get the name of the dish completely wrong.

The authors found that existing tests for these AI tools are almost entirely in English. They often judge the AI based on whether the picture looks good (visual plausibility), ignoring whether the words actually make sense (semantic correctness). It's like grading a student on how neatly they wrote a sentence, even if they spelled every word wrong.

2. The Solution: A New "Language Gym" (The Benchmark)

To fix this, the researchers built a massive training and testing ground called MULTITEXTEDIT.

  • The Setup: They took 300 base images (like blank posters) and created 12 different versions of each, one for 12 different languages (including English, Spanish, Arabic, Hebrew, Chinese, and even less common ones like Yoruba and Bengali).
  • The Control: Because every version started from the exact same image, they could isolate the language variable. It's like testing a car engine on a track where the only thing that changes is the fuel type, not the road or the weather.
  • The Scale: This resulted in 3,600 specific test cases covering 7 different types of edits (like changing font size, color, or replacing words).

3. The New Ruler: Measuring "Script Fidelity"

The researchers realized that standard computer tests (which just count how many pixels match) are terrible at spotting language errors.

  • The Analogy: Imagine a computer checking a handwritten note. If you write "Mother" but miss the accent mark on the "e" (making it "me" which means "tamarind" in Vietnamese), a pixel-counting robot might say, "Hey, that looks 99% identical!" But a human knows the meaning has completely changed.
  • The Fix: They created a new metric called Language Fidelity (LSF). They used a smart AI judge to look only at the specific words that were changed, checking for tiny details like missing accent marks, reversed letters (common in Arabic or Hebrew), or mixed-up scripts. This judge was trained to spot these "typos" that humans care about but machines often miss.

4. The Results: The "Global Layout" Trap

When they tested 12 different AI systems (both free and paid) on this new benchmark, they found a consistent pattern:

  • The "Good News": The AI is great at keeping the background looking nice. If you ask it to change text in a picture of a beach, the sand and ocean stay perfect.
  • The "Bad News": The AI struggles significantly with the actual words, especially in non-English languages.
    • The "Layout vs. Meaning" Mismatch: The AI often preserves the shape of the text box and the style of the font perfectly, but the letters themselves are gibberish, backwards, or missing critical marks.
    • The Hardest Languages: The AI stumbled the most on Hebrew and Arabic (which read right-to-left) and languages with complex accent marks.
    • The Easiest Languages: It did best on Dutch and Spanish, which are closer to English.

5. The Conclusion

The paper concludes that while AI image editors are getting better at "painting," they are still struggling with "writing" in languages other than English. They can copy the look of a foreign sign perfectly but often fail to get the words right.

The authors argue that to build truly useful tools for the whole world, we need to stop testing AI only in English and start measuring how well it handles the specific quirks of every writing system, from the direction of the letters to the tiny dots above them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →