← Latest papers
💬 NLP

Algospeak, Hiding in the Open: The Trade-off Between Legible Meaning and Detection Avoidance

This paper introduces a formal framework and dataset to analyze the trade-off between detectability and comprehensibility in "Algospeak" evasion strategies, defining a "Majority Understandable Modulation" threshold to quantify how linguistic alterations impact both human understanding and AI detection.

Original authors: Jan Fillies, Ronald E. Robertson, Jeffrey Hancock

Published 2026-05-08
📖 5 min read🧠 Deep dive

Original authors: Jan Fillies, Ronald E. Robertson, Jeffrey Hancock

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a giant, bustling town square. In this square, there are two main groups: the Town Criers (people or bots trying to spread messages) and the Town Guards (algorithms designed to stop harmful lies, scams, or toxic speech).

For a long time, the Town Criers just shouted their messages. But the Town Guards got smart; they learned to spot specific words and patterns that meant "danger." So, the Criers started playing a game of "hide and seek" with the words. They began swapping "vaccine" for "vaxx," "kill" for "k!ll," or using emojis to stand in for words. This new, modified way of speaking is called Algospeak.

This paper is like a scientific study of that game. The researchers wanted to understand the "Goldilocks Zone" of this hiding game: How much can you change your words to fool the guards without confusing the people you are trying to reach?

Here is a breakdown of their findings using simple analogies:

1. The "Too Much, Too Soon" Problem (The Trade-off)

The researchers discovered a delicate balance.

  • If you change your words too little: The Town Guards catch you immediately.
  • If you change your words too much: The Town Guards might miss you, but now your friends (the human audience) can't understand you either. You've turned your message into a secret code that only a few people know, defeating the purpose of spreading a message to a crowd.

They call the perfect middle point the "Majority Understandable Modulation" (MUM). Think of it like turning up the volume on a radio. If you turn it up just a little, the static (the guards) drowns out the music. If you turn it up too much, the music becomes so distorted that even your friends can't tell what song it is. The MUM is the exact volume knob setting where the music is clear enough for your friends to hear, but the static is just loud enough to confuse the machine.

2. The Experiment: A "Fake News" Lab

To test this, the researchers didn't use real people. Instead, they built a digital laboratory using AI robots (Large Language Models) to play the roles of both the Town Criers and the Town Guards.

  • The Setup: They took 20 common sentences about COVID-19 misinformation (like "Drinking bleach cures the virus").
  • The Distortion: They created 700 versions of these sentences. Some had just a few letters changed (like "c0vid"), some used abbreviations, some used emojis, and some used completely made-up words.
  • The Test: They asked different AI models two questions:
    1. The Guard Test: "Is this a lie?" (Can the AI detect the misinformation?)
    2. The Friend Test: "What did this sentence originally mean?" (Can the AI reconstruct the original meaning?)

3. What They Found

The results showed a clear pattern, like a slide going down:

  • As the words got more "scrambled" (more modulation): The AI "Guards" got worse at spotting the lies.
  • At the same time: The AI "Friends" got worse at understanding what the scrambled words meant.

The Surprising Twist:
The researchers found that different types of scrambling work differently.

  • The "Code Word" Strategy: Replacing a word with a secret code (like using "lemon" to mean "virus") was very good at fooling the guards, but it made the message hard to understand very quickly.
  • The "Spelling Change" Strategy: Just changing a letter to a number (like "v1rus") was easy for the guards to catch because it's a common trick they already know.
  • The "Phonetic" Strategy: Writing words how they sound (like "Kovit") was surprisingly hard for the AI to understand, even though it looked simple to humans.

4. The "Red Team" Warning

The authors are very careful to say they are not teaching bad guys how to win. Instead, they are acting as "Red Teamers" (like security testers who try to break a bank vault so the bank can fix the lock).

They are saying: "We found exactly where the lock is weak. If you are building a security system (a content moderator), you need to know that if people start using 'Code Words,' your current system will fail. But if you make the system too strict, you might accidentally block normal people who are just trying to talk."

Summary

This paper is a map of the "Hiding in Plain Sight" game. It proves that there is a specific limit to how much you can distort language to fool a computer before you lose your human audience. It gives us a way to measure that limit so we can build better tools to spot bad actors without silencing normal conversation.

Important Note: The study was a "proof of concept" using AI and synthetic (fake) examples. It did not test real humans or real-world social media platforms, so these are the first steps in understanding the problem, not the final solution.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →