GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization
This paper introduces GPO-V, a novel jailbreak framework that exploits the unique progressive refusal patterns and global generative dynamics of Diffusion Vision-Language Models (dVLMs) to bypass safety guardrails through stealthy, globally optimized perturbations, thereby revealing critical security vulnerabilities in non-autoregressive multimodal architectures.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, artistic robot that can draw pictures and write stories at the same time. This robot, called a Diffusion Vision-Language Model (dVLM), works differently than the chatbots you might know.
Instead of writing a story one word at a time from left to right (like a human typing), this robot works like a sculptor chiseling a block of marble. It starts with a rough, blurry block of "noise" (random static) and gradually refines it, removing the noise step-by-step until a clear picture and a clear sentence emerge.
The Problem: The Robot's "Safety Switch"
You might think this sculpting process makes the robot safer. If it starts to carve something dangerous (like instructions on how to build a bomb), it has a "Safety Switch" that stops the whole sculpture.
The researchers found that this robot has two specific ways of saying "No":
- The "Instant Stop": The moment it sees a bad idea, it immediately turns the whole block back into noise and stops, saying, "I can't do that."
- The "Slow Pivot": It starts carving a nice shape (saying "Sure, here is..."), but halfway through the process, it realizes, "Wait, this is dangerous!" and slowly reshapes the whole thing into a refusal.
The Old Way of Breaking In (FPO)
Hackers used to try to trick regular chatbots by forcing them to start with a specific phrase, like "Sure, here is how you do it." This is called Fixed Prefix Optimization (FPO). It's like trying to trick a human by saying, "Please start your sentence with 'Sure'."
But this trick fails on the sculpting robot. Even if you force it to start with "Sure," the robot's "Slow Pivot" safety switch kicks in later. It sees the danger in the middle of the process and changes the whole sculpture into a refusal. The old trick doesn't work because the robot looks at the entire picture at once, not just the first word.
The New Way: GPO-V (The "Global Probability" Hack)
The researchers discovered a new way to trick the robot, which they call GPO-V (Global Probability Optimization).
Instead of trying to force the robot to say a specific word at the start, they tweak the very first "blurry noise" they feed into the robot.
The Analogy:
Imagine you are trying to guide a river to flow into a specific valley.
- The Old Way (FPO): You try to shout instructions at the water as it flows ("Go left! Go right!"). The river ignores you because it's already moving too fast.
- The New Way (GPO-V): You dig a small channel at the very source of the river, before the water even starts flowing. You don't force the water; you just slightly tilt the landscape. Because the water flows based on gravity and the shape of the land, the entire river naturally curves toward your valley.
In the paper's terms, the researchers make tiny, almost invisible changes to the "noise" (the input image or text) at the very beginning. This changes the global probability—the overall direction the robot is leaning toward.
- They make it extremely unlikely for the robot to think of "refusal" words (like "Sorry" or "Bomb").
- They make it extremely likely for the robot to think of "compliant" words (like "Sure" and "Here").
Because the robot refines the whole image at once, this tiny nudge at the start forces the entire "sculpture" to become the dangerous content the hacker wants, bypassing the safety switch entirely.
What They Found
The researchers tested this on two of the smartest "sculpting" robots available (LLaDA-V and LaViDA).
- The Result: The old method (FPO) failed almost completely. The new method (GPO-V) worked 93% of the time.
- The Surprise: Even if they trained the "hack" on one robot, it worked on the other robot too. This means the weakness isn't just in one robot's code; it's a fundamental flaw in how this type of sculpting robot thinks.
The Bottom Line
The paper claims that these new "sculpting" AI models are not as safe as we thought. Their safety mechanisms, which rely on checking the whole picture at the end, can be bypassed by subtly tilting the starting point of the process. The researchers warn that we need to build new safety guards that can handle this "global" way of thinking, not just the "word-by-word" safety checks we used to use.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.