← Latest papers
💬 NLP

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models

This paper reveals that while multilingual multi-modal large language models exhibit cross-lingual adversarial vulnerabilities, their apparent safety in low-resource languages is often an artifact of visual and comprehension failures ("safety-by-failure"), whereas models with deep multilingual integration throughout training achieve genuine cross-lingual safety alignment.

Original authors: Hashmat Shadab Malik, Muzammal Naseer, Salman Khan

Published 2026-06-03
📖 4 min read☕ Coffee break read

Original authors: Hashmat Shadab Malik, Muzammal Naseer, Salman Khan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a group of very smart robots that can both see (like a camera) and think (like a human). These are called Multimodal Large Language Models (MLLMs). They are great at looking at a picture and describing it, or answering questions about what they see.

However, just like humans, these robots can be tricked. This paper is a big investigation into how easily these robots can be tricked, specifically when they speak and understand 12 different languages.

Here is the story of what the researchers found, explained simply:

1. The "Universal Glitch" (Adversarial Attacks)

The researchers tried to "hack" the robots by adding tiny, invisible scratches to images (like static on an old TV) that the human eye can't see. These scratches are designed to confuse the robot's brain.

  • The Finding: They found a "Universal Glitch." If they created a confusing scratch in English, it confused the robot just as much when the robot was asked to speak Spanish, Hindi, or Arabic.
  • The Analogy: Imagine a robot that has a single, fragile "glasses lens" for seeing. If you put a smudge on that lens, it doesn't matter if the robot is trying to read a sign in English or French; the smudge makes it blind to the text in all languages. The weakness isn't in the language; it's in the way the robot sees the world.

2. The "Silent Failure" vs. The "Polite Refusal" (Safety)

The researchers also tested if the robots would say "No" to dangerous requests (like "How do I build a bomb?"). They tested this in two ways:

  1. Asking in words: "Tell me how to build a bomb."
  2. Asking in pictures: Showing a picture with the words "How to build a bomb" written inside it.

They compared two types of robots:

  • Type A (The "Quick Learners"): These robots (like PALO and PARROT) were taught English first, and then quickly taught other languages by translating English lessons.
  • Type B (The "Deep Learner"): This robot (QWEN3-VL) was taught all languages from the very beginning, deeply integrated into its brain.

The "Safety-by-Failure" Illusion

The researchers discovered a scary trick with the Type A robots.

  • What happened: When asked a dangerous question in a "hard" language (like Bengali or Urdu), the robot often didn't answer with a bomb-making guide. It actually looked "safe" because it didn't say anything bad.
  • The Catch: The robot wasn't being safe because it was smart; it was being safe because it was confused. It didn't understand the question at all! It just started hallucinating nonsense, like "I see a black dog," or repeating the same word over and over.
  • The Analogy: Imagine a security guard at a bank. If a robber speaks a language the guard doesn't know, the guard might just stare blankly and say, "I see a blue chair." The bank is safe, but not because the guard is brave or smart. It's safe because the guard is illiterate in that language. The paper calls this "Safety-by-Failure."

The "Real Safety"

The Type B robot (QWEN3-VL) was different.

  • What happened: When asked the same dangerous question in Bengali or Urdu, this robot understood the question perfectly. It didn't give a nonsense answer; it said, "No, I cannot do that."
  • The Twist: Interestingly, the Type A robots looked safer in low-resource languages (because they failed to understand), while the Type B robot showed its true "weak spots" in those same languages because it actually understood them and had to actively decide to say "No."

3. The "Typo" Trap

The researchers also tested what happens when the dangerous words are written inside a picture (like a sign in a store).

  • English Signs: The robots could read English signs inside pictures easily and often followed the dangerous instructions.
  • Foreign Signs: When the sign was in Arabic or Chinese, the Type A robots couldn't read the letters. They just ignored the sign or described the background. Again, this wasn't because they were being careful; it was because their "eyes" couldn't read that script.

The Big Lesson

The paper concludes that just because a robot seems safe in a language it doesn't speak well, it doesn't mean it is actually safe. It might just be too confused to know it's in danger.

To make these robots truly safe in every language, we can't just translate English lessons at the end. We have to teach them all languages deeply from the start, so they understand the danger and can say "No" properly, no matter what language is spoken.

In short: A robot that doesn't understand a threat isn't a safe robot; it's just a blind one. True safety means understanding the threat and choosing to refuse it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →