InsertFuse: A Unified Framework for Multi-Category Reference-Guided Image Insertion
InsertFuse is a unified framework for multi-category reference-guided image insertion that achieves state-of-the-art performance by decoupling category-specific expert training from consolidation via Insertion On-Policy Distillation, while enhancing spatial control and reference fidelity through Token-Aligned Geometry Conditioning, Region-Balanced Flow Matching, and Reference CFG.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a master chef who can cook anything from a delicate soufflé to a hearty stew. Now, imagine someone asks you to make a single dish that is both a perfect soufflé and a perfect stew at the exact same time, using the same pot and the same set of instructions. You might try, but the result would likely be a soggy mess: the heat needed for the stew would ruin the delicate rise of the soufflé, and the flavors would clash. This is the kind of problem computer scientists face when teaching artificial intelligence (AI) to edit photos.
In the world of AI, "diffusion models" are like these master chefs. They are incredibly smart systems that can create or change images by slowly turning random static (like TV snow) into a clear picture, step by step. Recently, these models have gotten so good that we can ask them to "put a cat in this chair" or "change the sky to sunset." But there's a catch: different types of edits require very different skills. Putting a tiny ring on a finger needs fine, delicate precision. Putting a whole person into a scene needs to understand how their body bends and how their clothes look. Trying to teach one single AI model to be an expert at all these different tasks at once is like asking that chef to cook the soufflé and the stew simultaneously; the AI gets confused, and the results often look weird or blurry.
This is where a new paper called InsertFuse comes in. The researchers, led by a team from Shanghai Jiao Tong University and others, realized that the old way of training these AI chefs was the problem. Instead of forcing one model to learn everything at once, they built a clever two-step system. First, they trained five different "expert" chefs, each specializing in just one type of ingredient: one for accessories (like jewelry), one for animals, one for clothes, one for general objects, and one for humans. Each expert became a master of their specific domain, learning exactly how to handle the unique quirks of their category without getting distracted by the others.
But having five different chefs is annoying if you just want one app to do everything. So, the team invented a special training technique called Insertion On-Policy Distillation (IOPD). Think of this as a "student" chef who watches the five experts work. The student doesn't just memorize the recipes; it practices making the dish itself, and whenever it gets stuck, it asks the correct expert for the exact move to make next. If the student is trying to put a hat on a head, it asks the "accessory expert." If it's trying to put a person in a room, it asks the "human expert." By learning from the right expert at the right moment, the student becomes a unified master chef who can handle any insertion task perfectly, without the confusion of trying to learn everything at once.
To make sure the student chef doesn't just guess where to put things, the team also added two special tools. The first is Token-Aligned Geometry Conditioning, which is like giving the AI a precise GPS map of exactly where the new object should go, ensuring it fits the shape of the hole perfectly without bleeding into the background. The second is Region-Balanced Flow Matching, which acts like a fair judge. In normal training, the AI might ignore a tiny object (like a bracelet) because the huge background (like a wall) takes up most of the screen. This new tool tells the AI, "Hey, even though the bracelet is small, it's super important, so pay extra attention to it," ensuring that tiny details aren't lost.
Finally, to make sure the inserted object looks exactly like the reference photo (keeping its unique texture and identity), they used a technique called Reference CFG. This is like giving the student chef a strict rule: "You must keep the original pattern on this shirt, no matter how you stretch it."
The results are impressive. When the researchers tested their new "student" model against other top AI editors, it won in almost every category. It preserved the look of the original objects better, placed them more accurately, and kept the background looking natural. In a user study where people voted for their favorite images, InsertFuse was chosen by over 50% of participants, far beating the competition. The paper suggests that by separating the learning of specific skills from the process of combining them, we can build AI that is both a specialist and a generalist, solving the "soggy stew" problem of image editing once and for all.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.