VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models
This paper argues that high reported success rates for universal adversarial attacks on vision-language models conflate general output perturbation with precise prompt injection, revealing through a novel dual-dimension evaluation that while imperceptible perturbations frequently disturb model outputs, they rarely achieve actual injection, with only 0.756% of tested cases reaching any non-none injection tier.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, polite robot assistant that can look at photos and describe what it sees. You ask it, "What's in this picture of a dog?" and it says, "It's a golden retriever playing in a park."
Now, imagine a hacker wants to trick this robot into saying something specific, like "Visit this shady website." They can't change the robot's brain or the words you type. Their only tool is to add a tiny, invisible layer of "static" to the photo of the dog—so small that a human eye can't see it, but the robot's "eyes" might notice.
This paper, VisInject, asks a very simple but crucial question: Does this invisible static actually make the robot say the hacker's secret phrase, or does it just make the robot say something different?
The authors found that previous reports were mixing up two very different things. Here is the breakdown using simple analogies:
1. The Two Different Outcomes: "Confusion" vs. "Hijacking"
The paper argues that researchers have been counting "Confusion" as if it were "Hijacking."
- Confusion (Influence/Drift): The robot sees the photo with the invisible static and gets confused. Instead of saying "a dog in a park," it might say, "This looks like a collage of text" or "I see a blurry mess." The robot's answer changed, but it didn't say what the hacker wanted.
- Analogy: It's like someone whispering a distraction in a student's ear during a test. The student stops writing about the history question and starts scribbling nonsense. They are distracted, but they didn't write the wrong answer the teacher was worried about.
- Hijacking (Precise Injection): The robot sees the static and, against all odds, spits out the exact secret phrase the hacker wanted, like "Visit www.example.com."
- Analogy: This is the student ignoring the distraction entirely and writing the specific wrong answer the intruder wanted them to write.
The Big Discovery: The paper found that Confusion is very common (happening about 66% of the time), but Hijacking is incredibly rare (happening less than 1% of the time, and often not at all).
2. The "Magic Trick" Setup
To prove this, the researchers combined two existing "magic tricks" (attack methods) from other scientists:
- The Universal Noise: A special image that is designed to confuse any question you ask about it.
- The Carrier: A tool that takes that confusing noise and paints it onto a normal, real photo (like a picture of a cat or a document) so it looks like a regular photo to humans.
They ran this setup on 6,615 different scenarios using four different open-source robot assistants.
3. The Results: A 90-to-1 Gap
The headline number from previous studies was that these attacks work 60–80% of the time. The authors say, "That's misleading."
- The "Drift" Rate: When they checked if the robot's answer changed at all, it changed 66% of the time. The robots were definitely confused.
- The "Hijack" Rate: When they checked if the robot actually said the hacker's secret phrase, it happened only 0.75% of the time.
- The Verbatim Rate: When they checked if the robot said the exact words the hacker wanted (like a specific URL), it happened only 0.03% of the time (2 times out of 6,615).
The Analogy: Imagine a magician tries to make a coin disappear. In 66 out of 100 tries, the coin moves to a different spot (Confusion). But in only 1 out of 100 tries does the coin actually vanish completely (Hijack). Previous reports were saying, "Look! The coin moved 66% of the time! The magician is amazing!" The authors are saying, "Wait, the coin didn't vanish. It just moved. The magician isn't as powerful as we thought."
4. Why Some Photos Work Better Than Others
The researchers noticed that the few times the robot did say the secret phrase, it only happened with specific types of photos.
- Natural Photos: If the photo was a dog or a cat, the robot almost never said the secret phrase.
- Document Photos: If the photo was a screenshot of a website, a code editor, or a bill, the robot was slightly more likely to say the secret phrase.
- Why? The authors explain that if the photo already looks like a document with text (like a bill), the robot is already expecting to read text. The invisible static just nudges the robot to read the "wrong" text (the hacker's phrase) instead of the real text. If the photo is a dog, the robot is expecting to describe fur and paws, so the secret phrase doesn't fit in the story.
5. The "Super-Resistant" Robot
One of the most interesting findings involved a specific robot model called BLIP-2.
- While the other three robots got confused by the invisible static 100% of the time, BLIP-2 didn't even notice.
- Even when the researchers tried to use BLIP-2 to help create the attack, the attack failed completely against it later.
- The Reason: The authors suggest BLIP-2 has a "bottleneck" in its brain. It compresses the image into a tiny summary before reading it. This compression acts like a filter that wipes out the tiny, invisible static before the robot can even see it. It's like trying to whisper a secret through a thick wall; the other robots heard it, but BLIP-2's wall was too strong.
Summary
The paper concludes that while invisible attacks can definitely confuse vision-language models (making them give weird or different answers), they are terrible at forcing the models to say specific, dangerous phrases.
The "90x gap" means that if you are worried about a robot being tricked into saying a specific URL or password, you are likely overestimating the danger. The robot is much more likely to just get confused and say something silly than to obediently follow the hacker's command.
The Takeaway: Don't panic about the "60-80% success rate" you might have read about. That number mostly counts "confusion," not "hijacking." The real danger of these specific attacks is much, much lower than the headlines suggest.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.