Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models
This paper reveals a critical vulnerability in cloud-edge Large Vision-Language Model (LVLM) inference by demonstrating that a black-box adversary can manipulate as little as 10% of transmitted vision tokens to reduce model accuracy by up to 88.31% across various state-of-the-art models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart robot assistant (a Large Vision-Language Model) that lives in a giant, powerful data center in the "Cloud." You, however, are on a small, battery-powered device like a smartphone or a smart camera at the "Edge."
To get the robot to answer your questions about an image you took, you don't send the whole heavy image to the cloud. Instead, your phone does a quick, lightweight scan and turns the image into a set of digital "notes" called Vision Tokens. It then sends these notes over the internet to the cloud, where the big brain finishes the job and sends back an answer.
The New Weak Spot: The Messenger
The paper argues that this "hand-off" between your phone and the cloud creates a new, dangerous weak spot. Think of the internet connection as a messenger running between you and the robot.
The researchers discovered that a hacker (an adversary) standing in the middle of this path can intercept those digital notes. They don't need to break into the robot's brain or hack your phone; they just need to tweak the notes while they are being carried.
The Attack: "Vision Token Manipulation"
The researchers call this a VTM-Attack (Vision Token Manipulation Attack). They tested four different ways a hacker could mess with these notes, even if they could only change a tiny fraction of them (just 10%):
- The Shuffle (Permutation): Imagine the notes are a deck of cards describing a picture. The hacker swaps the order of a few cards. The robot still sees the same cards, but the story they tell is now jumbled and confusing.
- The Eraser (Masking): The hacker takes a few notes and turns them into blank paper (zeros). It's like looking at a photo where a few crucial pieces of the puzzle are suddenly missing.
- The Static (Gaussian Perturbation): The hacker adds a little bit of "static" or noise to the notes, making them slightly fuzzy or distorted, like a radio signal with interference.
- The Flip (Sign Flip): This was the most destructive. Imagine a note says "Red." The hacker flips it to mean "Not Red" or the exact opposite direction in the robot's mind. It's like taking a map and flipping the North arrow to point South. The robot gets completely turned around.
The Results: A House of Cards
The team tested this on six of the smartest vision robots available today (ranging from small to massive models). The results were shocking:
- Tiny changes, huge crashes: By changing just 10% of the notes, they could crash the robot's intelligence.
- The "Flip" was deadly: In some cases, the "Sign Flip" attack reduced the robot's accuracy from nearly 88% down to almost 0%. It's as if a smart person suddenly forgot how to read or see after a single, subtle whisper in their ear.
- Not all robots are equal: Some robot models (like the Qwen family) collapsed almost instantly under these attacks. Others (like the InternVL family) were a bit tougher, but still vulnerable.
The "Smart" Hack
The researchers also found a way to make the attack even worse. Instead of randomly picking which notes to mess with, they used a mathematical method to find the most important notes (the ones the robot relies on most) and targeted those specifically. This made the robot fail even faster.
The Bottom Line
The paper concludes that while splitting the work between your phone and the cloud is efficient, it opens a door for hackers to sabotage the system by tampering with the data in transit. Even a small, targeted tampering with the "notes" sent from the edge to the cloud can cause the most advanced AI vision systems to fail completely.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.