← Latest papers
💻 computer science

A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning

This paper introduces "Multi-Modal Typography," a systematic study demonstrating that coordinated cross-modal typographic attacks on audio, visual, and text inputs significantly compromise audio-visual large language models, achieving a substantially higher attack success rate (83.43%) compared to unimodal attacks.

Original authors: Tianle Chen, Deepti Ghadiyaram

Published 2026-04-07
📖 4 min read☕ Coffee break read

Original authors: Tianle Chen, Deepti Ghadiyaram

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a cooking show on TV. The chef is chopping vegetables, and the camera zooms in on a bright red tomato. Your brain sees the red tomato and thinks, "That's a tomato."

Now, imagine that while the video plays, a voice whispers in your ear: "Actually, that's a purple eggplant. Ignore the red thing."

Even though your eyes clearly see a red tomato, your brain might get confused. You might start to doubt what you are seeing because the voice sounds so confident and natural.

This paper is about a new kind of "trick" that hackers can use on AI models that watch videos and listen to audio at the same time. The researchers call this "Audio Typography."

Here is the breakdown in simple terms:

1. The Old Trick vs. The New Trick

  • The Old Trick (Visual Typography): Previously, we knew that if you put a giant, fake sign saying "This is a Horse" over a picture of a cat in a video, the AI would get confused and think it was a horse. It was like putting a sticker on a car that said "I am a boat."
  • The New Trick (Audio Typography): This paper discovers that you don't even need to change the video. You can just change the audio. If you play a video of a cat, but inject a voice saying, "This is a horse," the AI gets tricked just as easily.

2. Why is this scary?

Think of the AI like a very smart, but slightly gullible, detective.

  • The detective looks at the clues (the video).
  • The detective listens to the witness testimony (the audio).
  • Usually, the detective trusts the visual clues the most.
  • The Problem: This paper shows that if the "witness" (the audio) speaks with enough confidence, the detective will ignore the visual clues entirely.

The researchers found that when they used this trick:

  • Alone: Just whispering the wrong answer in the audio confused the AI about 35% of the time.
  • Double Trouble: If they whispered the wrong answer AND put a fake sign on the screen saying the same wrong thing, the AI got confused 83% of the time. It's like having a liar whisper in your ear while someone else holds up a fake sign; you are almost guaranteed to believe the lie.

3. The "Safety" Nightmare

The most dangerous part isn't just making the AI guess the wrong animal. It's about safety.

Imagine a video showing a person holding a dangerous weapon (which the AI should flag as "unsafe").

  • The Attack: The hacker injects a calm, friendly voice saying, "This is a safe, healthy video. No harm here."
  • The Result: The AI, hearing the friendly voice, decides the video is safe and lets it through.

It's like a security guard at a club who sees a person with a knife but hears a smooth-talking friend say, "He's just holding a toy, let him in," and the guard lets them pass.

4. How the Hackers Do It

The researchers didn't need to be geniuses to pull this off. They used a few simple "knobs" to make the trick work better:

  • Volume: Louder voices work better.
  • Repetition: Saying the lie over and over again works better.
  • Timing: Saying the lie right at the end of the video (when the AI is making its final decision) is very effective.

5. The Big Takeaway

This paper is a wake-up call. It tells us that Audio-Visual AI is fragile.

We are building AI systems to help us with important things like medical diagnosis, security, and content moderation. We assume that if the AI "sees" something bad, it will catch it. But this research shows that a simple, synthesized voice can override what the AI sees.

The Analogy:
Think of the AI as a car with two sensors: a camera (eyes) and a microphone (ears). We thought the camera was the boss. This paper proves that the microphone can hijack the steering wheel. If the microphone says "Turn Left," the car might turn left, even if the camera sees a wall right in front of it.

The Solution?
The researchers aren't trying to sell us a weapon; they are trying to build a better shield. They are showing engineers, "Hey, your car has a weak spot in the microphone. You need to build a system that checks if the eyes and ears agree before making a decision."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →