← Latest papers
🤖 machine learning

ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks

The paper introduces ImagenWorld, a comprehensive benchmark comprising 3,600 condition sets and 20,000 fine-grained human annotations across six tasks and domains, designed to stress-test image generation models through explainable, localized error analysis and revealing that while models excel in artistic and photorealistic generation, they struggle significantly with editing and text-heavy domains.

Original authors: Samin Mahdizadeh Sani, Max Ku, Nima Jamali, Matina Mahdizadeh Sani, Paria Khoshtab, Wei-Chieh Sun, Parnian Fazel, Zhi Rui Tam, Thomas Chong, Edisy Kin Wai Chan, Donald Wai Tong Tsang, Chiao-Wei Hsu, T
Published 2026-03-31
📖 5 min read🧠 Deep dive

Original authors: Samin Mahdizadeh Sani, Max Ku, Nima Jamali, Matina Mahdizadeh Sani, Paria Khoshtab, Wei-Chieh Sun, Parnian Fazel, Zhi Rui Tam, Thomas Chong, Edisy Kin Wai Chan, Donald Wai Tong Tsang, Chiao-Wei Hsu, Ting Wai Lam, Ho Yin Sam Ng, Chiafeng Chu, Chak-Wing Mak, Keming Wu, Hiu Tung Wong, Yik Chun Ho, Chi Ruan, Zhuofeng Li, I-Sheng Fang, Shih-Ying Yeh, Ho Kei Cheng, Ping Nie, Wenhu Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you've just built a fleet of incredibly talented digital artists. Some can paint a sunset from a description, others can take an old photo and swap the dog for a cat, and some can even combine three different photos into one new masterpiece. These are the Image Generation Models (like DALL-E, Midjourney, or the new ones mentioned in the paper).

But here's the problem: How do you know if they are actually good?

Until now, testing these artists was like grading a painting exam with a broken ruler. Some tests only checked if the picture looked "pretty" (ignoring if it made sense), others only checked if the text matched the picture, and many gave a single score without telling you why the artist failed. It was like getting a "C-" on a math test without seeing which numbers you got wrong.

Enter ImagenWorld. Think of this as the ultimate, stress-test driving course for these digital artists.

🚗 The Driving Test (The Benchmark)

The researchers created a massive "driving test" called ImagenWorld. Instead of just asking the cars to drive in a straight line, they threw every possible scenario at them:

  • The Terrain: They tested the models on six different "roads": Art, Real-life photos, Charts/Graphs, Text-heavy posters, Computer graphics, and Screenshots (like a picture of a website).
  • The Maneuvers: They didn't just ask for a new picture. They asked the models to:
    • Paint a new scene from scratch.
    • Take an existing photo and change just the sky.
    • Use one reference photo to change the style of another.
    • Use three different reference photos to build a complex new image.

In total, they created 3,600 unique challenges (like "Draw a cat wearing a hat" or "Change the text on this receipt to say 'Free'").

👮 The Judges (Human vs. Robot)

To grade the results, they used two types of judges:

  1. Human Experts: Real people looked at the images and gave them scores. But they didn't just give a number. They used a special "X-ray vision" tool to tag exactly what went wrong.
    • Example: "The text is gibberish," or "The cat has six legs," or "The shadow is on the wrong side."
  2. AI Judges (VLMs): They also asked a super-smart AI to grade the images to see if it could replace the humans.

🔍 What Did They Discover? (The Plot Twists)

After running the tests on 14 different models, the researchers found some surprising things:

1. The "Undo" Button Problem (Editing is Harder)
It's much easier for these models to paint a new picture from scratch than to edit an existing one.

  • The Analogy: Imagine asking a chef to cook a new meal vs. asking them to take a finished lasagna and just change the cheese. The chefs often either:
    • Option A: Throw away the whole lasagna and cook a brand new one (ignoring your request to just change the cheese).
    • Option B: Hand you the exact same lasagna back, pretending they did nothing.
    • The Lesson: Models struggle to make small, precise changes without breaking the whole image.

2. The "Text Trouble" (Charts and Screenshots are Tough)
Models are amazing at painting beautiful sunsets or portraits. But if you ask them to edit a screenshot of a website or fix a bar chart, they often fail miserably.

  • The Analogy: It's like a painter who can draw a realistic horse but gets confused when asked to fix a typo on a menu. The numbers in the charts often don't add up, and the text turns into alien symbols.
  • The Exception: One model, Qwen-Image, was surprisingly good at this. Why? Because its creators specifically fed it a diet of "text-heavy" images during training. It proves that data matters just as much as the model's brain.

3. The "Closed-Source" vs. "Open-Source" Gap
The models owned by big tech companies (Closed-Source, like GPT-Image-1) generally won the race. They were the most consistent drivers.

  • The Analogy: Think of the big companies as having a massive, well-funded racing team with the best fuel. The open-source community (the "DIY" racers) is catching up fast in some areas (like painting new pictures), but they still struggle with the complex editing maneuvers.

4. Can AI Judge AI?
The researchers asked: "Can we stop using humans and just let AI grade the art?"

  • The Verdict: Sort of, but not quite. The AI judges were great at ranking which image was "better" overall (about 79% accurate compared to humans).
  • The Catch: The AI judges couldn't explain why an image was bad. They couldn't point to the extra leg on the cat or the broken chart. For that level of detail, humans are still irreplaceable.

🏁 The Takeaway

ImagenWorld isn't just a list of scores; it's a diagnostic tool. It tells us exactly where these digital artists are stumbling.

  • They are getting better at making pretty pictures.
  • They are still terrible at making precise edits.
  • They struggle with text and numbers.
  • And while AI can help grade the work, we still need human eyes to tell us exactly what went wrong so we can teach the models how to fix it.

This paper is a roadmap for the future: to build image generators that don't just look good, but actually understand our instructions perfectly, whether we want a new painting or a tiny tweak to an old photo.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →