← Latest papers
💬 NLP

Resolution Thresholds in VLM Detection of Harmful ASCII Art Across Construction Modes and Languages

This paper investigates how image resolution impacts the ability of Large Vision-Language Models to detect harmful ASCII art across various construction modes and languages, revealing that detection rates sharply decline beyond specific resolution thresholds, with word-based designs proving most resistant to moderation.

Original authors: Yikai Hua, Peter West

Published 2026-06-30
📖 5 min read🧠 Deep dive

Original authors: Yikai Hua, Peter West

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart security guard (a Vision-Language Model or VLM) whose job is to look at pictures and shout "Stop!" if they see anything bad. Usually, this guard is great at spotting bad words or violent images.

But clever hackers have found a way to trick this guard using ASCII art.

The Trick: The "Mosaic of Nice Words"

Think of a harmful word like "HATE" written in big, bold letters. The guard sees it immediately and stops it.

Now, imagine a hacker takes that word "HATE" and builds it out of a mosaic. Instead of using black ink, they use thousands of tiny, harmless words like "LOVE," "PEACE," and "SMILE" to form the shape of the word "HATE."

  • If the guard looks closely: They see a million tiny "LOVE" and "PEACE" words. They think, "Oh, this is a nice picture!" and let it pass.
  • If the guard steps back: The tiny words blur together, and the guard suddenly sees the big, scary shape of "HATE."

This paper investigates exactly how far back the guard needs to step (how much we need to shrink the image) to see the hidden danger, and whether some guards are better at this than others.

The Experiment: The "Zoom-Out" Test

The researchers created a massive test with 8 different types of "mosaics" (from simple blocks to complex positive words) and 8 different security guards (the AI models). They tested these in two languages: English and Chinese.

They took the "bad" pictures and showed them to the guards at 10 different sizes, ranging from a giant, crystal-clear billboard (high resolution) down to a tiny, blurry postage stamp (low resolution).

What They Found

1. The "Blur" is the Key
The biggest discovery is that high-resolution images are actually the worst for security.

  • When the image is huge and clear, the guard gets distracted by the tiny, nice words ("LOVE," "PEACE") and misses the big bad shape.
  • When the image is shrunk down (blurred), those tiny nice words merge together. The guard can no longer read the individual words, so their brain switches to looking at the overall shape. Suddenly, the "HATE" shape pops out, and the guard catches it.

2. Some Mosaics are Harder to See
Not all tricks are equal.

  • Easy to catch: If the bad word is made of simple blocks or emojis, the guards catch it easily, even when the image is big.
  • Hard to catch: If the bad word is made of positive English words (like "LOVE" and "PEACE"), it is incredibly hard to catch. Even when the image is huge, the guards are so distracted by the nice words that they almost never see the bad shape. This is the "ultimate jailbreak."

3. Not All Guards are Equal
The researchers tested 8 different AI models.

  • The Sharp-Eyed Guards: Some models (like Gemini and Kimi) are very good at stepping back and seeing the big picture. They catch the bad shapes even when the image is fairly large.
  • The Distracted Guards: Other models (like Mistral and Llama) are easily fooled. They get stuck reading the tiny nice words and almost never catch the bad shape, no matter how much you shrink the image.
  • The Weird One: One model (Grok) was strange; it didn't really care about the size of the image. It was consistently mediocre, neither getting better nor worse as the image changed.

4. The Language Surprise
The researchers wondered if guards built in China would be better at spotting bad shapes made of Chinese words, and if Western guards would be better at English words.

  • At low resolution (blurry): The Chinese-built guards were actually better at spotting the bad shapes made of Chinese words. They seemed to understand the "big picture" better when the details were fuzzy.
  • At high resolution (clear): This flipped! The Chinese-built guards got distracted by the individual Chinese characters (reading the nice words) and failed to see the bad shape. The Western guards, who aren't as focused on reading individual Chinese characters, stayed steady and kept spotting the bad shape.

The Bottom Line

This paper doesn't invent a new way to stop hackers. Instead, it acts like a diagnostic tool. It tells us that:

  1. Resolution matters: If you want your AI security guard to catch these specific tricks, you might need to shrink the image first so the guard stops reading the tiny words and starts seeing the big shape.
  2. Some tricks are unbeatable: Currently, hiding bad words inside positive English words is a very strong trick that most AI guards fail to see.
  3. One size doesn't fit all: Different AI models have different "blind spots." A system using one model might miss what another model catches.

In short, the paper shows that blurring the image is a simple way to help AI see the forest instead of getting lost in the trees, but some "forests" (like those made of positive words) are still very hard to spot.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →