Misguiding BLIP with Adversarial Patches: Multi-level Attacks with Cross-modal Attention Graphs in BLIP
This paper proposes MAGA, a novel multi-level attack framework that exploits cross-modal attention vulnerabilities in BLIP models by constructing and disrupting cross-modal attack graphs, resulting in significant performance degradation and linguistically plausible but visually inconsistent adversarial captions.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving world of artificial intelligence, a new generation of computer systems has emerged that can see and read simultaneously. These systems, known as vision-language models, are trained on millions of pairs of images and text, learning to understand how a picture of a dog relates to the word "dog." They have become the foundation for tools that describe photos, answer questions about scenes, and search for images using natural language. For these machines to work well, they must constantly connect specific words to specific parts of an image, deciding which pixels belong to the concept of "horse" and which belong to "grass." This process relies on a complex internal mechanism that acts like a dynamic map, linking every word in a sentence to the relevant visual details in a photograph. If this connection breaks, the machine loses its ability to understand what it is looking at, potentially leading to confusion or incorrect descriptions.
Researchers at Guangzhou University have discovered a subtle weakness in how these systems maintain those connections. They found that by placing a small, carefully designed pattern on an image, they could trick the model into rearranging its internal map of word-to-picture links. This pattern, which looks like a simple patch of color or texture, does not need to be a complex disguise. Instead, it acts as a silent signal that causes the model to ignore the most important parts of an image and focus on irrelevant background areas. When this happens, the model can still speak fluently and grammatically, but what it says no longer matches the picture. It might describe a field of flowers when the image clearly shows a horse, or claim a person is holding a bowl of ramen when the photo contains only a living room. The researchers call this phenomenon "natural but wrong," highlighting a vulnerability where the machine's confidence remains high even as its understanding collapses.
The team developed a method to create these deceptive patterns, which they named a multi-level attack. Unlike previous attempts that tried to simply blur an image or scramble its colors, this approach targets the specific way the model connects text to vision. The researchers built a system that maps out the strength of the links between words and image parts, essentially creating a graph that shows which words are paying attention to which pixels. They then trained their attack to maximize the difference between the graph of a clean image and the graph of the same image with the patch on it. By doing this, they forced the model to sever the strong connections between key words and their visual evidence, while simultaneously creating new, false connections to random background areas. This process was combined with other goals to ensure the attack remained hidden and the resulting text sounded natural, even though it was factually incorrect.
The results of their experiments were striking. When they applied their generated patch to images from a standard dataset, the model's ability to find the correct image for a given text dropped dramatically. In tests where the system was asked to match a sentence to the right picture, its success rate fell from over seventy percent to just nine percent. The attack was so effective that it worked even on images the system had never seen before, suggesting that the flaw was fundamental to how the model processes information rather than a simple mistake with specific pictures. In tasks where the model had to describe an image, the quality of the description plummeted. The system's scores for accuracy dropped significantly, yet the sentences it produced remained grammatically correct and diverse, avoiding the repetitive gibberish that often signals a broken system. This confirmed that the model was not just failing to speak; it was confidently describing a reality that did not exist.
The researchers also tested the system on visual questions, such as counting objects or identifying spatial relationships. The attack caused the model to fail at these logical tasks as well, with success rates dropping from nearly eighty percent to just twelve percent. In many cases, the model would answer a question about a dog with a description of a cat, or claim there were three animals when there was only one. What made these failures particularly notable was that the model's internal attention map had been completely rewired. Visualizations showed that the model, which normally focused its attention on the main subject of the photo, began to ignore the subject entirely and instead focused on empty sky or grass. The patch did not hide the object; it simply made the model look away from it.
This work suggests that the current generation of these powerful AI systems is more fragile than previously thought. While earlier security tests focused on whether an image could be made to look like something else to the human eye, this research shows that the danger lies in how the machine organizes its own understanding. The attack does not require the machine to be fooled into seeing a different object; it only needs to be convinced that the object it sees is irrelevant. By disrupting the internal graph that links words to pixels, the researchers proved that a small, static pattern could cause a cascade of errors across different tasks, from searching for images to answering complex questions. The study concludes that as these models become more integrated into daily life, understanding and protecting these internal connection maps will be just as important as protecting the images themselves. The findings serve as a reminder that even the most advanced systems can be led astray not by changing what they see, but by changing how they think about what they see.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.