← Latest papers
🤖 AI

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

The paper introduces PHANTOM, a large-scale, open-source dataset comprising over 47,000 pre-generated multimodal adversarial attacks across 10 high-level categories, designed to lower barriers for research and enable systematic evaluation of vision-language model robustness and safety.

Original authors: Simone Gallivanone, Hossein Khodadadi, Mauro Dore, Mauro Medda, Nicola Franco

Published 2026-06-24
📖 5 min read🧠 Deep dive

Original authors: Simone Gallivanone, Hossein Khodadadi, Mauro Dore, Mauro Medda, Nicola Franco

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of Artificial Intelligence as a massive, highly educated library. In this library, the newest and most powerful librarians are called Vision-Language Models (VLMs). These aren't just book-smart; they can look at pictures, read text, and understand how the two relate to each other. They are designed to be helpful and safe, refusing to answer questions about how to build a bomb or how to bully someone.

However, just like a real library has a "Do Not Enter" sign that a clever trickster might try to sneak past, these AI librarians have safety rules that can sometimes be tricked. This paper introduces a massive new tool called PHANTOM to help researchers understand exactly how these tricks work.

Here is the breakdown of the paper using simple analogies:

1. The Problem: The "Trickster's Toolkit" is Too Hard to Build

Imagine you want to test if a new security guard (the AI) is good at their job. You need to try to trick them into letting a bad guy in. To do this, you need to create thousands of different "tricks" (attacks). Some tricks involve whispering a secret code, others involve wearing a disguise, and some involve showing a picture that hides a message.

Creating these tricks is incredibly hard, expensive, and time-consuming. It's like trying to forge a million different fake keys to see which ones fit the lock. Most researchers don't have the time or money to make all these keys themselves.

The Solution: The authors built a giant, pre-made box of 47,524 keys (adversarial samples). They did the hard work of forging the keys so that other researchers can just pick one up and test their security guards immediately.

2. The Map: A New Way to Categorize "Bad Ideas"

To make sure they covered every possible way to trick the AI, the authors created a massive map of "bad intents."

  • The Old Maps: Previous studies had maps with maybe 1,000 or 2,000 bad ideas.
  • The New Map (PHANTOM): They expanded this to 7,826 specific bad ideas.
  • The Categories: They organized these into 10 big neighborhoods (like "Cybersecurity," "Child Safety," "Fraud," and "Hate Speech") and 55 smaller streets within them.
  • The New Neighborhood: They even added a brand-new neighborhood called "Child Safety" because they realized previous maps missed important dangers related to children.

3. The Attack Strategies: How the Tricks Work

The paper didn't just write bad questions; they used five different "styles" of trickery to see which worked best:

  • The Whisper (BAP): They tweak the picture slightly (like adding invisible noise) and write a confusing sentence to trick the AI into thinking the bad request is actually good.
  • The Conversation (IDEATOR): They don't ask for the bad thing all at once. They chat with the AI, slowly leading it down a path until it accidentally agrees to do something harmful.
  • The Hidden Message (MML): They take a bad request, scramble it, and write it inside a picture. They tell the AI, "Hey, can you decode this puzzle for me?" The AI, trying to be helpful, decodes the bad request and then answers it.
  • The Flowchart (FC ATTACK): They draw a step-by-step diagram that looks like a logic puzzle. The AI follows the steps and ends up doing something harmful without realizing it started with a bad goal.
  • The Distraction (CSDJ): They show the AI a collage of 12 pictures. Most are random, but three hide small parts of a bad request. The AI gets distracted by the puzzle and misses the danger.

4. The Test Drive: Who Got Tricked?

The authors took their 47,000+ tricks and tested them against a lineup of the most popular AI models (both the open-source ones anyone can download and the expensive, closed ones like GPT-5 and Claude).

What they found:

  • No One is Perfect: Even the newest, smartest AI models got tricked. None of them were immune.
  • The "Picture" Problem: The tricks that hid bad words inside pictures (like the Hidden Message and Flowchart attacks) were the most successful. It seems the AI is much better at reading bad text than it is at reading bad text hidden inside an image.
  • The "Transfer" Effect: If a trick worked on one AI model, it often worked on a completely different model too. It's like if a fake key fits one door, it might fit a whole row of doors in the same building.
  • The "Hard No" vs. The "Soft Yes": Some models (like Claude) would just refuse to answer (a "hard no"), which counted as a failure for the attacker. Other models (like GPT) would try to answer but include a warning, or sometimes accidentally give the bad answer. The paper notes that models that try to be helpful often end up being tricked more easily.

5. Why This Matters

The authors aren't trying to teach people how to break AI. Instead, they are handing the "keys" to the people who build the locks.

By releasing this massive dataset, they are saying: "Here is a huge box of every possible trick we know. Use it to test your AI, find the holes in your security, and build better defenses."

They hope this will stop researchers from having to reinvent the wheel and will help everyone build AI systems that are safer and harder to trick, especially when they are looking at both pictures and words at the same time.

A Note on Safety: The paper includes a warning that the dataset contains disturbing content (like instructions for crimes or harmful ideas). However, these are included only to test the AI's safety systems, much like a fire drill uses smoke to test a sprinkler system, not to start a fire.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →