← Latest papers
💻 computer science

MDTD-ArtIR: Benchmarking Image Editing and Restoration Models for Art Image Restoration under Texture-Overlay Degradations

This paper introduces the MDTD-Art dataset and benchmark to evaluate image restoration under texture-overlay degradations, revealing that image editing models enhanced by structured prompt engineering outperform specialized restoration architectures in preserving semantic and structural details.

Original authors: Mridula Vijendran, Shuang Chen, Hubert P. H. Shum

Published 2026-08-04
📖 6 min read🧠 Deep dive

Original authors: Mridula Vijendran, Shuang Chen, Hubert P. H. Shum

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are an art restorer, but instead of a dusty brush and a palette of paints, you have a super-smart computer program. Your job is to fix old, damaged paintings—maybe one that has cracked like a dry riverbed, another stained with mysterious yellow spots, or a third where a chunk of the canvas has simply vanished. This is the world of AI-driven art restoration, a field where computers try to guess what a masterpiece looked like before time and accidents ruined it. But here's the tricky part: the computer doesn't just need to fill in the blank spots; it needs to do it without making up wild, fake details that never existed. It's like trying to finish a friend's story when they stop talking halfway through; you want to guess the ending that fits their style, not invent a dragon if they were writing a mystery.

To test if these computer programs are actually good at this, scientists usually give them practice problems. But most of these practice problems are too easy or too fake. They might ask the computer to fix a blurry photo or remove a few raindrops, which is like practicing on a slightly smudged window before trying to fix a shattered stained-glass window. Real art damage is messy, weird, and often hides parts of the picture behind layers of grime or cracks. The big question researchers are asking is: Can our best AI tools actually handle this messy, real-world damage without losing the soul of the artwork?


The Great Art Heist: When AI Tries to Fix a Painting

Meet the MDTD-Art team, a group of researchers from Durham University who decided to build a new, super-challenging training ground for AI art restorers. They realized that the old practice tests were like playing "Simon Says" with a broken record—the instructions were too simple, and the damage was too predictable. So, they invented a new game called MDTD-ArtIR.

Think of their new dataset as a "damage simulator." Instead of just blurring a picture, they took beautiful, clean paintings and overlaid them with 47 different types of weird textures—like cracked mud, splattered paint, or porous sponges. But they didn't just slap these textures on; they played with a "transparency knob." They could make the damage faint (low opacity), like a light dusting of flour, or make it thick and heavy (high opacity), like pouring a bucket of mud over the painting. This created a spectrum of difficulty, from "easy peasy" to "good luck, you'll need a miracle."

The Contest: Specialized Mechanics vs. Creative Storytellers

The researchers set up a showdown between two types of AI models:

  1. The Specialized Mechanics (Universal Image Restoration models): These are the "fix-it" bots. They are trained specifically to remove scratches, noise, or blur. They are like a master watchmaker who knows exactly how to fix a broken gear.
  2. The Creative Storytellers (Image Editing Models): These are the "make-it-up" bots. They are usually used to change a photo (like turning a dog into a cat or adding a hat). They are like a creative writer who can improvise a whole new scene based on a few clues.

The team wanted to see who could better restore the paintings when the damage was heavy and the clues were scarce. They tested these models using two different "instruction manuals" (prompts):

  • The Simple Prompt: Just said, "Fix this."
  • The Expert Prompt: Said, "You are an expert restorer. The mask is damage, not art. Remove it and rebuild the original scene."

The Big Surprise: The Storytellers Won (Mostly)

Here is the twist that the paper discovered: The Creative Storytellers (Image Editing Models) generally outperformed the Specialized Mechanics.

When the damage was light (low opacity), everyone did okay. But as the "mud" got thicker (high opacity), the Specialized Mechanics started to struggle. They would often leave the damage there or try to fix it in a way that looked stiff and unnatural. They were too focused on the "rules" of fixing and didn't have enough imagination to guess what was underneath.

The Creative Storytellers, however, were much better at guessing the missing parts. When the researchers gave them the Expert Prompt (the one that clearly explained the task), these models got even better. For example, the Nano Banana (NB) Pro model jumped from a score of 9.63 dB to 11.87 dB in image quality when given the expert instructions at high opacity. The Flux 2 model saw a similar jump, going from 8.57 dB to 11.01 dB. It's as if the storyteller suddenly understood the assignment: "Oh, I'm not supposed to paint over the mud; I'm supposed to imagine what was under it!"

The Catch: When Imagination Goes Too Far

But there's a catch. Because these "Storyteller" models are so good at making things up, they sometimes get too creative. In the worst cases, they didn't just fix the painting; they changed the story entirely.

The paper shows some funny (but sad) examples where the AI looked at a cracked face and decided, "I think this person should have a different nose," or looked at a landscape and decided to add a mountain that wasn't there. This is called "hallucination." The FireRed model, which is trained mostly on faces, got so confused by face-shaped cracks that it "hijacked" the image and turned the whole painting into a weird, blurry face. The GPT Image 1.5 and NB Pro models were the best at keeping the original identity, but even they sometimes changed the colors or the shape of the picture slightly.

The Verdict: Clarity is Key

The main takeaway from this paper is that for fixing really damaged art, how you ask the question matters more than the tool you use.

  • Specialized tools (like AutoDIR) are great at keeping the original picture safe but often fail to remove heavy, weird textures.
  • Creative tools (like NB Pro and Flux) are amazing at guessing the missing parts, but they need clear instructions to stop them from making up fake details.
  • The "Expert Prompt" is the secret sauce. When the AI is told clearly, "This is damage, not art," it performs much better, especially when the damage is heavy.

The authors suggest that while these creative models are powerful, they aren't perfect yet. They can still get confused and change the identity of the artwork. But by using the right instructions and understanding that these models are better at "guessing the story" than "fixing the gears," we can get much closer to restoring our precious art without losing its soul. It's a reminder that in the world of AI, sometimes the best way to fix a broken picture is to ask the computer to imagine what it should look like, rather than just telling it how to patch the holes.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →