← Latest papers
🤖 AI

HP-Edit: A Human-Preference Post-Training Framework for Image Editing

The paper introduces HP-Edit, a post-training framework that utilizes a human-preference-aligned automatic evaluator (HP-Scorer) and a new real-world dataset (RealPref-50K) to effectively apply Reinforcement Learning from Human Feedback to diffusion-based image editing, significantly improving model alignment with human preferences.

Original authors: Fan Li, Chonghuinan Wang, Lina Lei, Yuping Qiu, Jiaqi Xu, Jiaxiu Jiang, Xinran Qin, Zhikai Chen, Fenglong Song, Zhixin Wang, Renjing Pei, Wangmeng Zuo

Published 2026-04-22
📖 5 min read🧠 Deep dive

Original authors: Fan Li, Chonghuinan Wang, Lina Lei, Yuping Qiu, Jiaqi Xu, Jiaxiu Jiang, Xinran Qin, Zhikai Chen, Fenglong Song, Zhixin Wang, Renjing Pei, Wangmeng Zuo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very talented but slightly stubborn digital artist. This artist (an AI model) is incredibly good at following instructions like "put a dog here" or "change the sky to sunset." However, sometimes the artist gets the spirit of the request wrong. They might make the dog look like a cartoon monster, or the sunset might look like a neon sign. The result is technically "correct" based on the words, but it feels fake or ugly to a human eye.

This paper introduces HP-Edit, a "coaching system" designed to teach this digital artist how to please human tastes, not just follow literal commands.

Here is the breakdown of how it works, using simple analogies:

1. The Problem: The "Robot" Artist

Current image-editing AI is like a student who studied hard for a test but only memorized the textbook. It knows how to swap a cow for a giraffe, but it might forget to make the giraffe look like it belongs in the photo.

  • The Issue: The AI was trained on massive amounts of data, but that data is a mix of cartoons, fake images, and real photos. It doesn't know what humans actually find beautiful or realistic.
  • The Result: The edits often look "off"—weird shadows, blurry edges, or objects that don't fit the scene.

2. The Solution: HP-Edit (The "Human Taste Coach")

The authors created a three-step training camp to fix this.

Step A: Hiring a "Taste Judge" (The HP-Scorer)

Before you can teach the artist, you need a way to grade their work. Humans are great at this, but they are slow and expensive to hire for millions of images.

  • The Innovation: The team trained a super-smart AI (called a Visual Large Language Model) to act as a Taste Judge.
  • How it works: They showed this Judge thousands of examples of "good" and "bad" edits and gave it a specific checklist (e.g., "Does the new object look like it belongs? Are the shadows natural?").
  • The Metaphor: Think of this Judge as a strict food critic. Instead of just saying "this burger is cooked," they say, "The bun is too dry, and the patty is burnt." This Judge learns to give a score from 0 to 5, just like a human would.

Step B: Finding the "Hard Cases" (RealPref-50K)

If you only show the artist easy examples (like "add a red ball to a white wall"), they won't learn much. They need to practice on the tricky stuff.

  • The Innovation: The team built a massive dataset called RealPref-50K containing 50,000 real-world editing scenarios.
  • The Filter: They used their "Taste Judge" to scan through thousands of images and throw away the easy ones. They kept only the "Hard Cases"—the edits where the AI usually fails or where the result is barely acceptable.
  • The Metaphor: Imagine a math teacher who throws away all the easy 1+11+1 problems and only gives the student the hardest calculus problems to solve. This forces the student to actually learn how to think.

Step C: The "Reinforcement Training" (RL Post-Training)

Now, the artist (the AI model) goes back to school with the new "Hard Case" dataset and the "Taste Judge."

  • How it works: The AI tries to edit an image. The Taste Judge gives it a score. If the score is low, the AI gets a "punishment" (it learns not to do that). If the score is high, it gets a "reward."
  • The Result: The AI learns to tweak its internal settings to maximize the Taste Judge's score. Since the Judge mimics human preference, the AI starts producing results that look much more natural and realistic.

3. The Results: From "Okay" to "Wow"

The paper tested this new "coached" AI against the best existing models.

  • Before HP-Edit: The AI might swap a person for a dog, but the dog looks like a floating plastic toy.
  • After HP-Edit: The AI swaps the person for a dog that blends perfectly into the lighting, shadows, and texture of the photo. It looks like a real photo taken by a human photographer.

Summary Analogy

Think of the original AI as a new chef who knows how to chop vegetables and boil water (the technical skills) but has never tasted food (human preference).

  • HP-Edit brings in a Master Chef (the HP-Scorer) to taste every dish.
  • Instead of letting the new chef practice on easy toast, they force them to practice on complex, tricky dishes (the RealPref-50K dataset).
  • The Master Chef gives constant feedback: "Too salty," "Burnt," "Needs more spice."
  • Eventually, the new chef learns not just how to cook, but how to cook delicious food that people actually want to eat.

In short: HP-Edit bridges the gap between "the AI did what I said" and "the AI did what I wanted," making digital image editing feel much more human and magical.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →