← Latest papers
💻 computer science

Detection and Safeguarding of Chinese Toxic Content from Users and LLMs: A Survey

This paper presents a comprehensive survey that systematically reviews research on detecting user-generated toxic content on Chinese social media and safeguarding Chinese large language models against LLM-generated toxicity, aiming to consolidate fragmented progress and identify key challenges for future development.

Original authors: JunYu Lu, Weiming Wang, Deyi Ji, Lanyun Zhu, Bo Xu, Liang Yang, Hongfei Lin, Roy Ka-Wei Lee

Published 2026-08-24
📖 5 min read🧠 Deep dive

Original authors: JunYu Lu, Weiming Wang, Deyi Ji, Lanyun Zhu, Bo Xu, Liang Yang, Hongfei Lin, Roy Ka-Wei Lee

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The internet is a vast, bustling public square where billions of people share thoughts, jokes, and stories every day. But like any crowded place, it has a shadow side: toxic content. This includes insults, hate speech, bullying, and subtle mockery that can hurt individuals and poison the atmosphere of online communities. For years, researchers have built tools to spot this harmful language, mostly focusing on English, where the rules of insult and offense are well-mapped. However, the Chinese-speaking world presents a unique challenge. In Chinese online culture, people often hide their true intent behind a veil of slang, homophones (words that sound the same but are written differently), and cultural references that only insiders understand. A phrase might look innocent to a computer but carry a vicious punch to a human reader. Furthermore, a new layer of complexity has emerged with the rise of large language models. These powerful artificial intelligence systems can now generate text that mimics human conversation, but they can also be tricked into producing harmful content or used by bad actors to create it at scale.

A team of researchers from several Chinese universities has now taken a comprehensive look at how the field is handling these specific challenges. They conducted a systematic review, which is essentially a deep dive into all the existing studies, datasets, and methods related to detecting and stopping toxic content in Chinese. Their work covers two main fronts: first, how to identify harmful content created by human users on social media, and second, how to protect large language models from generating or being manipulated into producing such content. The researchers found that while progress has been made, the field is still fragmented. There is no single, unified way to define what counts as toxic in Chinese, and the tools currently available often struggle with the dynamic, evolving nature of online slang and the subtle, implicit ways people express hostility.

The researchers began by mapping out the landscape of data. To teach computers to recognize toxicity, scientists need massive collections of examples. They found that researchers have gathered data from major Chinese platforms like Weibo and Zhihu, creating datasets that cover everything from direct insults to more complex forms of hate speech and sarcasm. However, they noted a significant problem: the criteria for labeling something as toxic are often inconsistent. One study might label a comment as harmless banter, while another marks the same comment as hateful, depending on the cultural context or the specific group being targeted. This lack of agreement makes it difficult to compare different tools or to build a system that works reliably across different situations. The review also highlighted that while there are many datasets for text, there is a growing need for resources that include images, videos, and memes, as toxicity often hides in the combination of text and visuals.

On the technical side, the team examined how researchers are trying to solve these detection problems. Early methods relied on simple keyword matching, but these fail when users swap characters for homophones or use internet slang to bypass filters. To counter this, newer approaches try to understand the deeper meaning and cultural context of a message. Some methods use "knowledge enhancement," feeding the computer extra information about cultural references or regional slang to help it understand the hidden intent. Others use "multimodal" techniques, which allow the system to look at both the text and the accompanying image or video to catch the full picture of the insult. The researchers also reviewed how these systems are trained, noting that combining different tasks—like learning to detect sarcasm at the same time as detecting hate speech—can help the models become more robust. However, they found that many of these advanced methods still struggle when faced with entirely new types of attacks or rapidly changing slang.

The second half of the review focused on the new frontier: safeguarding large language models. These AI systems are designed to be helpful, but they can be "jailbroken"—tricked by cleverly worded prompts into ignoring their safety rules and generating toxic content. The researchers surveyed the various ways people are testing these models, from simple questions to complex, multi-turn conversations designed to wear down the AI's defenses. They found that while some models have become better at refusing harmful requests, they are not perfect. A major issue is "over-sensitivity," where the AI refuses to answer harmless questions because it mistakes them for something dangerous, such as refusing to discuss certain historical figures or cultural topics. The review also pointed out that the methods used to train these models to be safe are often a black box; we know they work to some degree, but we don't fully understand why they fail in specific cases or how to fix them without breaking other capabilities.

The authors conclude that while the field has moved forward, significant hurdles remain. The most pressing challenge is the dynamic nature of Chinese internet culture; as soon as a detection tool learns to spot a specific type of insult, users invent a new way to say it. The researchers suggest that future work needs to focus on creating more consistent standards for what counts as toxic, developing systems that can learn and adapt quickly to new slang, and improving the fairness of these tools so they don't unfairly censor specific groups or dialects. They also emphasize the need for better tools to understand the complex interplay between text, images, and video. Ultimately, the goal is to build systems that can protect online spaces without stifling free expression, a balance that requires a deep understanding of both the technology and the human culture it serves. The review serves as a roadmap, showing where the current tools fall short and pointing the way toward more resilient and culturally aware solutions for the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →