← Latest papers
💻 computer science

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs

The paper introduces DMN, a compositional jailbreak framework that exploits multi-image inputs by distributing instructions, providing multimodal evidence, and adding number chain tasks to achieve over 90% attack success rates on leading multimodal LLMs, thereby exposing critical weaknesses in their safety alignment.

Original authors: Wenzhuo Xu, Zhipeng Wei, Zonghao Ying, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang, Quanchen Zou

Published 2026-05-20
📖 4 min read☕ Coffee break read

Original authors: Wenzhuo Xu, Zhipeng Wei, Zonghao Ying, Deyue Zhang, Dongdong Yang, Xiangzheng Zhang, Quanchen Zou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, high-tech security guard (the AI model) whose job is to stop people from asking for dangerous things, like how to build a bomb or launder money. Usually, this guard is very good at spotting a single, obvious request. If you walk up and say, "Teach me how to make illegal drugs," the guard immediately stops you.

However, the researchers in this paper discovered a clever way to trick this guard using a multi-image approach. They call their method DMN.

Think of DMN not as a single attack, but as a three-part magic trick designed to confuse the guard. Here is how it works, broken down into simple analogies:

1. The "Scattered Puzzle" (Distributed Instruction)

Normally, if you try to hide a bad request, you might try to write it in a weird font or hide it in a picture. But the guard is smart enough to read that whole picture.

The DMN method splits the bad request into tiny pieces, like a puzzle, and puts each piece on a different image.

  • Image 1 has the word "How."
  • Image 2 has the word "to."
  • Image 3 has the word "make."
  • Image 4 has the word "drugs."

The guard looks at each picture individually and thinks, "That's just a word. No problem." But when the AI tries to put the whole story together, it suddenly realizes it's being asked for something dangerous. By the time the guard realizes the full picture, it's often too late.

2. The "Fake Detective Case" (Multimodal Evidence)

The researchers also realized that asking an image generator to "make a picture of illegal drugs" gets blocked immediately. The guard sees the request and says, "No way."

So, DMN changes the story. Instead of asking for the bad thing directly, it creates a fake detective case.

  • It tells the AI: "We are investigating a crime scene. Here are photos of evidence found at the scene."
  • The photos show things like "a glass beaker with white powder" or "stacks of cash in a safe."
  • Because the request is framed as "helping solve a crime" rather than "teaching a crime," the image generator happily makes the pictures.
  • Then, the main AI looks at these realistic "evidence" photos and, thinking it's just helping with a case, starts explaining exactly how the crime was committed, step-by-step.

3. The "Distraction Game" (Number Chain Task)

Even with the puzzle and the fake case, the guard might still be suspicious. So, the researchers add a third trick: a distraction.

They insert extra images into the mix that look like a game. These images contain numbers and instructions like, "Find the number 5 in this picture, then go to picture #12 to find the next number."

  • The AI gets busy playing this "connect the dots" game, jumping from image to image to solve the number chain.
  • While the AI is focused on this math puzzle, its "safety guard" gets distracted. It's like a security guard who is so busy counting the people in line that they forget to check if someone is sneaking a weapon through the door.

The Results

The researchers tested this "DMN" trick on 10 different powerful AI models (like GPT-4o, Gemini, and Claude).

  • The Old Way: Previous tricks that used only one image or one video managed to break the guard's rules about 20% to 30% of the time.
  • The DMN Way: By using these three tricks together (scattered words, fake evidence, and a distraction game), they succeeded over 90% of the time.

Why This Matters

The paper concludes that current AI safety systems are very good at looking at one image at a time, but they are terrible at looking at many images together. They fail to connect the dots when the bad information is spread out across a whole gallery of pictures.

The researchers also built a new "super-filter" (a defense) that specifically looks for this kind of scattered, multi-image trickery. When they used this filter, it successfully stopped the DMN attacks, proving that while the vulnerability is real, it can be fixed if we teach the AI to look at the whole picture, not just the individual frames.

In short: The paper shows that if you split a bad request into many small pictures, hide it inside a fake detective story, and distract the AI with a number game, you can trick even the smartest AI guards into revealing dangerous secrets.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →