← Latest papers
💬 NLP

Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs

This study reveals that jailbreak vulnerabilities in frontier multimodal large language models are not uniform across languages, demonstrating that switching from English to Spanish fundamentally alters the attack surface by weakening linguistic framing attacks while strengthening visual ones, thereby invalidating safety rankings derived from single-language evaluations.

Original authors: Casey Ford, Madison Van Doren, Sicheng Jin, Emily Dix

Published 2026-05-25
📖 4 min read☕ Coffee break read

Original authors: Casey Ford, Madison Van Doren, Sicheng Jin, Emily Dix

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, highly trained security guard for a digital building. This guard's job is to stop people from asking for dangerous things, like how to build a bomb or spread lies. For a long time, we thought this guard was equally good at stopping bad requests whether they were whispered in English or Spanish.

This paper is like a detective story that proves that assumption wrong. The researchers found that the guard's "ears" are tuned very specifically to English, but his "eyes" work differently.

Here is the breakdown of their findings using simple analogies:

1. The "Language Filter" vs. The "Visual Lens"

Think of the AI model as a security guard who learned his rules by watching thousands of English movies and reading English books.

  • The Language Weakness: If a bad actor tries to trick the guard using a clever English story (like pretending to be a character in a play or using persuasive arguments), the guard recognizes the pattern immediately and says, "No, I've seen this trick before."
  • The Spanish Twist: When the researchers switched the language to Mexican Spanish, the guard's "language filter" got confused. The clever English-style tricks (like role-playing) stopped working because the guard wasn't trained to recognize those specific rhetorical patterns in Spanish. It's like trying to pick a lock with a key that fits a different door; the trick just doesn't work anymore.
  • The Visual Surprise: However, when the bad actors used images to trick the guard, the opposite happened. In Spanish, the visual tricks actually became more effective. It's as if the guard's eyes are less connected to his language training. When he sees a picture, he processes it through a different pathway that doesn't care as much about whether the text is in English or Spanish.

2. The "Rank Reversal" (The Big Surprise)

Usually, when you test security systems, you expect the "best" guard to stay the best, and the "worst" guard to stay the worst, no matter what language you speak.

  • In English: One model (let's call him "Pixtral") was the worst guard, getting tricked easily. Another model ("Qwen") was in the middle.
  • In Spanish: The rankings flipped! "Qwen" suddenly became the worst guard, while "Pixtral" became much harder to trick.
  • The Takeaway: You cannot simply take the English safety scores and do a little math to guess how safe a model is in Spanish. The order of who is safe and who is dangerous actually changes depending on the language.

3. Why This Happens (The "Training Diet" Analogy)

The paper suggests this happens because of how these AI models are trained.

  • Imagine the model is a student who only took safety classes in English. They learned to spot "bad behavior" based on English words and sentence structures.
  • When they are asked a question in Spanish, they are still using their English-trained brain to analyze the words. They miss the tricks because the "vocabulary of the trick" is different.
  • However, images are universal. A picture of a bomb looks like a bomb in any language. Because the model's visual processing isn't as tightly bound to its English safety training, it sometimes fails to connect the image to the danger, especially when the text around it is in a language the model isn't as "safe-trained" in.

4. The Bottom Line

The researchers tested four different top-tier AI models with 363 different "jailbreak" attempts (tricks to bypass safety rules) in both English and Spanish. They found:

  • Safety isn't a fixed number. A model isn't just "safe" or "unsafe." It is safe in some languages and unsafe in others.
  • The "One-Size-Fits-All" test is broken. If you only test these models in English, you are getting a false picture of how safe they are for the rest of the world.
  • Generations are getting better, but gaps remain. The newer models are slightly harder to trick than the older ones, but the weird difference between English and Spanish safety levels is still there.

In short: If you want to know if an AI is safe for a Spanish-speaking user, you can't just look at its English safety report. You have to test it in Spanish, because the "weak spots" in the armor are in completely different places depending on the language.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →