← Latest papers
🤖 AI

Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection

This paper introduces ImageProtector, a user-side method that embeds nearly imperceptible visual perturbations into images to trigger refusal responses in Multi-Modal Large Language Models, effectively preventing the unauthorized extraction of sensitive information while demonstrating robustness against common countermeasures.

Original authors: Zedian Shao, Hongbin Liu, Yuepeng Hu, Neil Zhenqiang Gong

Published 2026-04-13
📖 4 min read☕ Coffee break read

Original authors: Zedian Shao, Hongbin Liu, Yuepeng Hu, Neil Zhenqiang Gong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a beautiful, personal photo you want to share on social media. You love the picture, but you're worried that a "digital detective" (an AI) might scan it, figure out exactly who you are, where you live, or what you were doing, and steal that private information.

This paper introduces a digital bodyguard for your photos called ImageProtector. Here is how it works, explained simply:

The Problem: The Overly Curious AI

Think of modern AI models (called MLLMs) as incredibly smart, but slightly nosy, librarians. If you show them a picture of a street, they can tell you the name of the street, the city, and even guess the time of day. While this is great for helping people, it's a nightmare for privacy. If a bad actor uses these AIs to scan millions of photos online, they could harvest private details about everyone at once.

The Solution: The "Invisible Shield"

The researchers created a tool called ImageProtector. Instead of blurring your face or putting a black bar over your photo (which ruins the picture), this tool adds a tiny, invisible layer of "static" to the image.

To your human eye, the photo looks exactly the same. It's like adding a single grain of sand to a beach; you can't see it, but it's there.

How It Works: The "Confused Robot" Trick

Here is the clever part. This invisible static isn't just noise; it's a secret command written in a language only the AI understands.

  1. The Setup: You take your photo and run it through ImageProtector. The tool adds this "secret command" to the pixels.
  2. The Attack: A bad actor downloads your photo and asks the AI, "Who is this person?" or "Where is this?"
  3. The Result: Instead of answering the question, the AI sees the secret command. It gets confused and thinks, "Oh no! This looks like a request I'm not allowed to answer. I must refuse!"
  4. The Refusal: The AI politely says, "I'm sorry, I can't help with that request," and stops analyzing the image.

It's like putting a "Do Not Disturb" sign on your front door that is invisible to humans but glows bright red to the mailman. The mailman (the AI) sees the sign and decides not to deliver the package (the private info), even though the door looks normal to everyone else.

Why It's Special

  • It's Universal: The researchers tested this on six different types of AI "librarians." No matter which one the bad guy uses, the secret command works.
  • It's Invisible: The changes to the image are so small that you can't see them, so your photo still looks great to your friends.
  • It's Proactive: You do this before you post the photo. You don't need to know who is going to attack you; you just protect the photo in advance.

Can the Bad Guys Fight Back?

The researchers also asked: "Can the bad guys just clean the image to remove this shield?"
They tried three common methods:

  1. Adding Random Noise: Like shaking the photo. This helped a little, but it also made the photo look grainy and ruined the AI's ability to answer any questions (even safe ones).
  2. Smoothing the Image: Like blurring it slightly. This removed the shield but also made the AI dumber and slower.
  3. Training the AI to Ignore It: Like teaching the AI to ignore the "Do Not Disturb" sign. This worked a bit, but it required a lot of computing power and still made the AI less accurate.

The Bottom Line

ImageProtector is a way for regular people to take control of their digital privacy. It turns a photo into a "trap" for nosy AIs. If a bad actor tries to analyze it, the AI trips over the invisible trap and refuses to give up any secrets. It's a simple, effective way to say, "My data is mine, and you can't read it."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →