← Latest papers
🤖 AI

Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives

The paper proposes Dualin, a two-stage inversion method that jointly recovers a human-interpretable text prompt and the exact latent noise of a target image to overcome the limitations of existing techniques and achieve state-of-the-art image fidelity with controllable editing capabilities.

Original authors: Xiaolong Liu, Junjian Li, Yuan Xiao, Jiaqi Deng, Dayong Ye, Tianqing Zhu, Huan Huo

Published 2026-07-30
📖 3 min read☕ Coffee break read

Original authors: Xiaolong Liu, Junjian Li, Yuan Xiao, Jiaqi Deng, Dayong Ye, Tianqing Zhu, Huan Huo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to reverse-engineer a magic recipe. You have a delicious cake in front of you, and your goal is to figure out exactly how to bake it again. In the world of computer science, specifically in a field called "Computer Vision," there are magical programs known as Text-to-Image Diffusion Models. These models are like incredibly talented chefs who can bake any cake you can describe, as long as you give them the right text instructions (prompts). But what if you want to go backward? What if you have a specific, perfect cake (an image) and you want to know exactly what text instructions were used to make it? This is called "prompt inversion."

For a long time, scientists tried to solve this by just looking at the cake and guessing the recipe. Some tried to tweak the recipe words mathematically until they matched, but the results were often messy, like a recipe that said "add flour!! and zardoz." Others tried to write clear, human-readable recipes, but when they tried to bake the cake again using those words, the result looked nothing like the original—it was missing the specific texture, lighting, and tiny details. The big question was: Why does the cake look different even when the recipe words seem right? The answer lies in a hidden ingredient that most people ignored: the "noise." Think of this noise not as garbage, but as the specific random spark that determines the cake's exact shape and structure. Without capturing that spark, you can't perfectly recreate the cake, no matter how good your recipe words are.

This paper introduces a new method called Dualin (Dual Inversion) that solves this puzzle by looking at two things at once: the recipe (the text prompt) and the hidden spark (the noise). The researchers argue that trying to reverse-engineer an image by only guessing the text is like trying to rebuild a house by only describing the paint color; you miss the blueprint. Their approach is a two-step dance. First, they use a team of smart AI tools—a Vision-Language Model to describe the scene, a CLIP system to find the right artistic keywords, and a Large Language Model to write a perfect, human-readable recipe. But they don't stop there. In the second step, they perform a special "noise inversion." This is like rewinding the baking process to find the exact random spark that created the cake's unique structure.

By combining the perfect recipe with the exact spark, Dualin can recreate the original image with stunning accuracy. The authors tested this on thousands of images and found that their method produces pictures that look almost identical to the originals, far better than previous methods. They also proved mathematically that this technique allows for precise editing. Because the "spark" (noise) holds the structure and the "recipe" (prompt) holds the meaning, you can change the recipe to swap a cat for a dog or change the background, and the new image will keep the exact same layout and lighting as the original. The paper suggests that this dual approach is the key to unlocking truly controllable and high-quality image editing, moving beyond just guessing the words to understanding the full magic of how these images are made.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →