CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions
This paper introduces CREval, an automated and interpretable QA-based evaluation pipeline, and CREval-Bench, a comprehensive benchmark with over 800 samples, to systematically assess and reveal the current limitations of state-of-the-art image editing models in handling complex and creative manipulation tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a magical art studio where you can tell a computer, "Turn this photo of my dog into a superhero wearing a cape made of clouds," and it tries to do exactly that. This is what Instruction-Based Image Editing is all about.
But here's the problem: How do you know if the computer actually did a good job? Did it listen to your instructions? Did it keep your dog looking like your dog? Is the picture blurry or weird?
Currently, checking these results is like trying to grade a student's essay by just looking at the final score without reading the teacher's notes. It's vague, and sometimes the "teacher" (the AI evaluator) is biased or doesn't explain why it gave a bad grade.
This paper introduces CREval, a new, super-smart way to grade these AI art projects. Think of it as hiring a detective instead of just a judge.
The Problem: The "Black Box" Judges
Before this paper, researchers used big AI models to look at an edited image and just say, "Score: 7/10."
- The Issue: We didn't know why it was a 7. Did the AI miss the cape? Did the dog look like a cat? The score was a "black box"—a mystery.
- The Gap: Existing tests were mostly for simple tasks like "remove the person" or "change the sky to blue." They weren't built for creative, wild ideas like "turn the warrior into a tropical fish plushie" or "make a comic strip about a cowboy horse."
The Solution: CREval (The Detective)
The authors created CREval, which stands for Creative Review Evaluation. Instead of asking the AI, "How good is this?", CREval asks the AI a series of specific Yes/No questions (like a quiz).
Imagine you are a teacher grading a student's drawing. Instead of just giving a grade, you ask:
- "Did the student draw the hat?" (Yes/No)
- "Is the hat blue, as requested?" (Yes/No)
- "Did they keep the student's original face?" (Yes/No)
- "Are the lines shaky and messy?" (Yes/No)
CREval does exactly this, but for AI images. It breaks the evaluation down into three main categories:
- Instruction Following (IF): Did the AI actually do what you asked? (e.g., "Did it turn the robot into a cake?")
- Visual Consistency (VC): Did the AI keep the important parts of the original photo? (e.g., "Is it still the same robot, just a cake robot, or did it turn into a totally different character?")
- Visual Quality (VQ): Does the picture look good? (e.g., "Are the hands weird? Is the texture blurry?")
The Playground: CREval-Bench
To test this new detective system, the authors built a giant playground called CREval-Bench.
- It contains over 800 complex, creative challenges.
- These aren't simple tasks. They are things like: "Turn this wedding photo into a traditional Chinese cartoon," or "Make this mountain look like a giant stone titan rising from the ground."
- It's like a "final exam" for AI art models, designed to be much harder than previous tests.
What Did They Find? (The Report Card)
They ran the top AI models (both free/open-source and paid/closed-source) through this new test. Here's the verdict:
- The Good News: The paid, "closed-source" models (like Seedream 4.0 and Gemini 2.5) are currently the best at following these wild instructions. They are like the top art students who can handle complex prompts.
- The Bad News: Even the best models struggle. They often fail to keep the original character's identity (Visual Consistency). For example, they might turn a person into a cake, but the cake looks nothing like the person.
- The Open-Source Hope: Free models are catching up fast! Models like Qwen-Image-Edit are doing surprisingly well, showing that the gap between free and paid tools is shrinking.
Why Does This Matter?
Think of CREval as a universal translator between human creativity and AI capability.
- For Developers: It tells them exactly where their AI is failing (e.g., "You're great at following instructions, but you keep messing up the faces").
- For Users: It gives us a reliable way to know which AI tool is actually good at handling our creative, weird, and wonderful ideas.
In short, CREval stops us from guessing if AI art is good. It gives us a clear, step-by-step report card that explains exactly what the AI did right and what it got wrong, helping us build better tools for creative expression.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.