← Latest papers
🤖 machine learning

Banana100: Breaking NR-IQA Metrics by 100 Iterative Image Replications with Nano Banana Pro

The paper introduces Banana100, a dataset of 28,000 images degraded through 100 iterative editing steps, to reveal that multi-turn image editing causes severe quality accumulation and that current no-reference image quality assessment metrics fail to detect this degradation, posing significant risks to the stability of agentic AI systems.

Original authors: Kenan Tang, Praveen Arunshankar, Andong Hua, Anthony Yang, Yao Qin

Published 2026-04-07
📖 4 min read☕ Coffee break read

Original authors: Kenan Tang, Praveen Arunshankar, Andong Hua, Anthony Yang, Yao Qin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a magical photocopier that doesn't just copy a picture; it listens to your voice commands to change the picture. You tell it, "Add an apple," and it does. You tell it, "Make it brighter," and it does. This is how modern AI image editors work.

The paper "Banana100" asks a scary question: What happens if you keep using this photocopier over and over again?

Here is the breakdown of their discovery, explained simply.

1. The "Xerox Effect" (Iterative Degradation)

Imagine you take a crisp, high-definition photo of a banana.

  • Step 1: You ask the AI to "copy this exactly." It makes a copy. It looks 99% perfect.
  • Step 10: You ask it to copy the new copy. It looks a little grainy, like an old TV with static.
  • Step 20: You keep going. The image is now covered in digital snow, the colors are weird, and the banana looks like a blurry blob.

The researchers found that AI models are terrible at "copying" their own work. Every time they touch an image, they introduce tiny, invisible errors. When you feed the result back into the machine for the next round, those tiny errors pile up like dust in a corner. After about 10 to 20 rounds, the image is a mess of noise and distortion.

2. The "Broken Thermostat" (The Evaluator Failure)

Here is where it gets really weird. Usually, when an image gets ugly, we have tools (metrics) to measure how bad it is. These tools act like a thermostat: if the room gets hot (the image gets bad), the thermostat should scream "Too hot!"

The researchers tested 21 of these "thermostats" (called NR-IQA metrics). They expected the tools to say, "Hey, this image is garbage now!"

Instead, the tools said, "Wow, this is beautiful!"

  • The Analogy: Imagine you take a pristine, clean white shirt and start rubbing it in mud. After 20 minutes, it's a brown, muddy mess. But your "Cleanliness Meter" suddenly beeps and says, "This is the cleanest shirt I've ever seen!"
  • The Reality: The AI tools were so confused by the specific type of "digital mud" the generator created that they thought the noise was actually a sign of high quality. In one famous example, a clean image got a "bad" score, but after 20 rounds of corruption, the same image got a "perfect" score.

3. The "Clueless Artist" (Instruction Failure)

Not only did the images get ugly, but the AI also started forgetting what it was supposed to do.

  • The Counting Fail: If you asked the AI to "add one apple" 10 times, it might add 10 apples in the first round, then forget to add any in the next, or suddenly add a whole orchard.
  • The Hallucination: The AI would look at its own messy, noisy output and write a report saying, "I did a great job! The details are perfect!" even though the image was covered in static. It was lying to itself.

4. The "Digital Snowball" (Why This Matters)

Why should we care?

  • The Training Loop: If AI models are trained on data created by other AI models, and that data is full of this invisible "digital snow," the next generation of AI will learn from garbage. This is called Model Collapse. It's like a student studying from a textbook that has been photocopied 100 times; eventually, the text becomes unreadable gibberish.
  • The Silent Poison: Because the "thermostats" (quality checkers) are broken, this garbage data is slipping into our databases. We are unknowingly feeding our future AI systems a diet of noise.

5. The Solution: Banana100

The authors created a dataset called Banana100.

  • They took 13 different images (a forest, a peacock, a plate of dumplings).
  • They asked an AI to edit them 100 times in a row.
  • They saved all 28,000 resulting images.

This dataset is like a "crash test dummy" for the AI world. It proves that current AI editors are fragile and current quality checkers are blind. By releasing this data, they hope scientists can build:

  1. Better Editors: AI that doesn't get messy when it copies itself.
  2. Better Checkers: Tools that can actually tell the difference between a clean image and a noisy one.

The Bottom Line

The paper warns us that while AI is amazing at making pictures, it is currently fragile when asked to keep editing them. It's like a game of "Telephone" played by a robot: the message (the image) gets distorted with every turn, and the robot's own quality control system is too broken to notice the distortion until it's too late.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →