VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation
This paper introduces VARM-Bench, a novel benchmark designed to evaluate the verifiable, structured reasoning of Chinese abusive speech moderation by requiring models to generate field-anchored rationales for six key decisions, thereby revealing that strong label-level performance often masks significant errors in complete moderation records.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The internet is a vast, noisy marketplace of human thought, where a single sentence can be a harmless joke, a sharp critique, or a genuine threat. For years, computer programs designed to police this space have focused on a simple question: is this post bad or not? They scan for angry words and flag them, much like a security guard checking a list of banned items. But this approach has a blind spot. A program might correctly label a sentence as "harmful" while completely misunderstanding why. It might think a person is insulting a friend when they are actually quoting an insult to criticize it, or it might miss that a comment attacking a specific group is actually a neutral observation about that group's behavior. In the complex landscape of Chinese social media, where tone, context, and hidden meanings shift the weight of words, getting the final label right is not enough. If the computer's reasoning is flawed, the decision is fragile, and the system cannot be trusted to handle the millions of subtle interactions that happen every day.
To solve this, researchers at Nanjing University of Aeronautics and Astronautics and other institutions have built a new testing ground called VARM-Bench. Instead of just asking a computer to say "yes" or "no" to a post, this new system forces the computer to explain its thinking in a structured way, like a judge writing a verdict that must cite specific evidence. The researchers created a dataset of 8,000 real comments from Chinese social media platforms, carefully selecting examples that are tricky to interpret. These include posts where someone quotes an insult to mock it, or where a criticism of a behavior is mistaken for an attack on a person's identity. For each comment, human experts wrote a "gold standard" explanation that identifies exactly who is being discussed, whether the author is attacking or opposing that person, and what specific type of harm is involved. This creates a complete record of the decision, not just a final score.
The researchers then asked various artificial intelligence models to read these comments and produce their own explanations, following a strict format that required them to identify six key pieces of information: the target of the comment, the type of target, how clearly it was mentioned, the author's stance, whether it was harmful, and the specific category of harm. The system then automatically checked if the AI's explanation matched the human experts' record. The results revealed a startling gap. Many models could get the final "harmful" or "not harmful" label correct, yet their reasoning was completely wrong. A model might correctly flag a post as harmful because it saw a specific swear word, but fail to realize that the word was being used in a quote to criticize the very behavior it described. In these cases, the model got the right answer for the wrong reason, a mistake that would be invisible in a standard test but dangerous in a real-world moderation system.
The study found that while some advanced models performed well when given specific instructions and examples, they still struggled with the most difficult cases, particularly those involving sarcasm, quoted speech, or indirect references. The researchers discovered that the biggest bottleneck was not in deciding if something was bad, but in correctly identifying who or what was being talked about. When the computer misidentified the target, the entire chain of reasoning collapsed, even if the final label happened to match the human answer. By forcing the models to lay out their reasoning step-by-step and checking every part of that logic, the new benchmark exposes these hidden errors. It shows that a system that cannot explain its decisions clearly is not truly safe, because it cannot distinguish between a genuine threat and a complex social interaction. This work provides a new, verifiable way to test whether AI moderation tools are actually understanding the language they are policing, rather than just memorizing patterns of offensive words.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.