← Latest papers
💻 computer science

Perceptual Gaps: ASCII Art and Overlapping Audio as CAPTCHA

This paper proposes and evaluates novel CAPTCHA methods using ASCII art and overlapping audio, demonstrating that current frontier large language models fail to solve these tasks effectively, thereby offering a promising but potentially temporary solution for distinguishing humans from bots.

Original authors: Choon-Hou Rafael Chong

Published 2026-04-07
📖 4 min read☕ Coffee break read

Original authors: Choon-Hou Rafael Chong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet is a giant, bustling party. For years, the bouncers (website owners) used a simple trick to keep out the troublemakers (bots): they asked guests to solve a puzzle, like "read this squiggly text" or "click on all the traffic lights." This was the CAPTCHA.

But recently, the troublemakers got a massive upgrade. They hired super-smart AI assistants (Large Language Models) that can read squiggly text and spot traffic lights faster than any human. The old bouncer tricks don't work anymore; the AI guests are walking right in.

This paper is about designing new, tougher bouncer tricks that rely on things humans have been good at for millions of years, but that AI still finds confusing or too expensive to do.

Here are the two new tricks the authors tested:

1. The "ASCII Art" Puzzle (The Visual Trick)

The Old Way: Showing a picture of a word with lines drawn through it.
The New Way: Showing a word made out of tiny computer symbols (like |, _, +, and #) arranged to look like letters. It's like a picture made of text.

  • The Analogy: Imagine a human looking at a cloud. Even though it's just a fluffy white shape, our brains instantly say, "That looks like a dragon!" We are wired to see patterns and fill in the gaps.
  • The Problem for AI: AI models are like a robot that only sees a list of ingredients. If you show it a cloud made of text, it gets confused. It sees a jumbled list of symbols (|, +, #) rather than a "dragon." It tries to read the symbols as words, not as a picture.
  • The Result: The authors tested the smartest AI models in the world (like GPT-5 and Gemini 3). When they showed these models ASCII art, the AI models got 0% correct. They couldn't even guess one letter right. Meanwhile, a human could solve it in a split second.

2. The "Cocktail Party" Puzzle (The Audio Trick)

The Old Way: Asking a bot to listen to a clear voice and type what it says.
The New Way: Playing a recording where two people are talking at the same time, or where there is loud background noise (like a busy cafe), and asking the bot to answer a specific question based on only one of the voices.

  • The Analogy: This is called the "Cocktail Party Effect." Imagine you are at a loud party. You can focus on your friend's voice and ignore the music and other conversations. Humans do this naturally.
  • The Problem for AI: AI is usually trained to hear everything. If you play it a noisy recording, it tries to transcribe every sound it hears, getting confused by the mix. It struggles to "tune out" the noise and focus on just one voice.
  • The Result: The AI models did okay when the audio was clear, but as soon as the authors added noise or overlapping voices, the AI's performance crashed. They started guessing randomly, while humans could still hear the answer clearly.

The "Too Expensive" Strategy

The authors also point out a third defense: Cost.

Even if AI eventually learns to solve these puzzles, it might be too expensive for the bad guys to use.

  • The Analogy: Imagine a toll booth. If it costs a human 1 second to pay the toll, it's fine. But if it costs a robot 100 seconds of computer power to figure out the puzzle, and the robot has to do this millions of times, it runs out of money (or electricity) very quickly.
  • The Goal: The goal isn't just to make the puzzle impossible for AI, but to make it so expensive for AI that it's not worth the trouble. It breaks the "economics" of the bot attack.

The Bottom Line

The paper concludes that:

  1. ASCII Art is currently a super-strong shield. AI is terrible at it, but humans find it easy.
  2. Noisy Audio is also a good shield, though it's harder to make and might be annoying for some people.
  3. The Future: We can't rely on making puzzles that are "impossible" for AI forever, because AI keeps getting smarter. Instead, we should rely on puzzles that are cheap for humans but too expensive for robots to solve at scale.

In short, the authors found a way to make the bouncer ask a question that the AI is too "dumb" (or too "rich" to afford) to answer, while the human guest just smiles and walks through.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →