← Latest papers
💬 NLP

Exploring the Capability Boundaries of LLMs in Mastering of Chinese Chouxiang Language

This paper introduces "Mouse," a specialized benchmark for evaluating Large Language Models on the Chinese internet subcultural language "Chouxiang," revealing that while current models excel at contextual semantic understanding, they struggle with other tasks due to specific limitations in mastering this evolving linguistic phenomenon.

Original authors: Dianqing Lin, Tian Lan, Jiali Zhu, Jiang Li, Wei Chen, Xu Liu, Aruukhan, Xiangdong Su, Hongxu Hou, Guanglai Gao

Published 2026-04-20
📖 5 min read🧠 Deep dive

Original authors: Dianqing Lin, Tian Lan, Jiali Zhu, Jiang Li, Wei Chen, Xu Liu, Aruukhan, Xiangdong Su, Hongxu Hou, Guanglai Gao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a giant, bustling city. In this city, most people speak "Standard Chinese," which is like the official language of the government and news broadcasts. But there's also a secret, underground neighborhood where the locals speak a very strange, coded dialect called Chouxiang Language (or "Abstract Language").

This paper is like a group of detectives (the researchers) who decided to test how well the city's most advanced AI detectives (Large Language Models, or LLMs) can understand this secret neighborhood. They built a special test called Mouse to see if the AI can crack the code.

Here is the breakdown of their investigation, explained simply:

1. What is "Chouxiang Language"?

Think of Chouxiang Language as a linguistic game of "Telephone" mixed with a scavenger hunt.

  • The Rules: Instead of saying "You are stupid," a user might write something that looks like gibberish, like "🧠👻" or "91 安排".
  • The Tricks:
    • Homophones (Sound-alikes): Using a word that sounds the same but means something else (like using a character for "Ning" to mean "You").
    • Visuals: Using emojis or breaking characters apart. A picture of a brain might mean "smart," and a skull emoji might mean "death" or "bones."
    • Secret Meanings: A phrase might look innocent but actually be a specific insult or a joke known only to people in that specific internet community.

Originally, this language was used to sneak around censorship (like writing a password to say something forbidden), but over time, it evolved into a fun, neutral way for young people to bond and make jokes.

2. The "Mouse" Benchmark (The Test)

The researchers created a test called Mouse (a pun on "Mouse" and "Mice," perhaps implying small but numerous tests). It's like a gym workout for AI, but instead of lifting weights, the AI has to solve puzzles involving this secret language.

The test has six different challenges:

  1. Translation: Can the AI translate the secret code back into normal Chinese? (e.g., turning "🧠👻" into "You are a smart ghost").
  2. Component Classification: Can the AI spot how the code was made? Did they use a sound-alike? A picture? A hidden meaning?
  3. Intent Recognition: Is the user being mean, making a joke, or just saying hello?
  4. Toxicity Detection: Is this actually a hate speech disguised as a joke, or is it just harmless fun?
  5. Meaning Selection: If you see a weird sentence, can you pick the right meaning from three choices?
  6. Cloze Completion: If you are in a conversation, can you guess the next weird sentence that fits perfectly?

3. The Results: The AI Got Lost

The researchers tested the world's smartest AI models (like GPT-5, Qwen, and DeepSeek) on this test. Here is what they found:

  • The "Smart" AI is actually quite dumb here: The top-tier AIs did terribly on most tasks. They struggled to decode the jokes or understand the hidden meanings.
  • The "Overthinker" Problem: Interestingly, the biggest, most powerful AI models sometimes did worse than slightly smaller ones. It's like a genius student who tries to solve a simple riddle by writing a 10-page essay, overthinking it until they get the wrong answer. The smaller models just guessed, and sometimes that worked better.
  • The "Toxic" Blind Spot: The AIs were actually better at spotting hate speech than they were at understanding the jokes. This is because most previous training focused on catching bad words, so the AI is hyper-sensitive to "bad stuff" but clueless about "funny stuff."
  • The Human Gap: When real humans who know the culture took the test, they scored nearly 100%. The AI, even the smartest one, couldn't come close. It's like trying to understand a complex inside joke from a group of friends when you weren't invited to the party.

4. Why Does This Matter?

The paper argues that AI is currently too Western-centric. It's great at understanding standard English or Chinese news, but it fails when it encounters the messy, creative, evolving slang of real internet culture.

  • The Metaphor: Imagine an AI that has read every dictionary in the world but has never hung out at a skate park or a gaming forum. It knows the rules of grammar, but it doesn't know the vibe.
  • The Danger: If we rely on AI to moderate internet content, it might delete harmless jokes because it thinks they are insults, or it might miss real threats because it thinks they are just "weird slang."

5. The Conclusion

The researchers built Mouse to show the AI community that culture matters. You can't just feed an AI more data; you need to teach it the context and the soul of how people actually talk online.

In short: The paper is a reality check. It says, "Hey AI, you might be smart, but you don't understand the internet's secret handshake yet. Here is a test to prove it, and here is how we can fix it."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →