← Latest papers
💬 NLP

Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework

This paper proposes an LLM-based multi-agent framework that personalizes content moderation by simulating individual user sensitivities through specialized agents, achieving a 32% accuracy improvement over traditional centralized systems while offering a scalable approach to balancing platform governance with user autonomy.

Original authors: Ewelina Gajewska, Michal Wawer, Katarzyna Budzynska, Jaroslaw A. Chudziak

Published 2026-05-06
📖 5 min read🧠 Deep dive

Original authors: Ewelina Gajewska, Michal Wawer, Katarzyna Budzynska, Jaroslaw A. Chudziak

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet as a massive, chaotic city square where millions of people are shouting, sharing stories, arguing, and posting memes all at once. Currently, the "security guards" (content moderation systems) trying to keep this square safe use a single, giant rulebook for everyone. They decide what is "harmful" based on a one-size-fits-all list. If a post breaks a rule, it gets removed. If it doesn't, it stays.

The problem, as the authors of this paper point out, is that people are different. What one person finds funny, another finds deeply hurtful. What one person sees as a bold political opinion, another sees as a personal attack. The current system ignores these individual feelings, often leaving people overwhelmed by content they find toxic, or conversely, removing things they actually enjoy.

This paper proposes a new way to run the security guard station. Instead of one rigid rulebook, they suggest a personalized, team-based approach called PRISM.

The Core Idea: A Personalized Security Team

Think of the PRISM system not as a single robot, but as a customized security team assigned to every single user. Here is how this team works, using simple analogies:

1. The "Sensitivity Profile" (Your Personal Rulebook)
Instead of starting with a blank slate, the system learns your specific "pain points."

  • The Analogy: Imagine you are very sensitive to loud noises but don't mind messy rooms. Your neighbor, however, hates mess but loves loud music. The PRISM system learns your specific thresholds. It builds a profile that says, "This user gets upset easily by insults, but they are fine with strong political arguments."
  • How it works: As you interact with the system (telling it, "I don't like this post" or "This is fine"), it updates your profile. It learns exactly what you find harmful, not what a generic algorithm thinks you should find harmful.

2. The "Manager Agent" (The Team Leader)
When a new post arrives, a central "Manager" looks at it and your profile.

  • The Analogy: The Manager is like a conductor in an orchestra. They look at the post and say, "This post has some aggressive language and mentions a specific group. We don't need the whole orchestra; let's just call in the Sociologist (who knows about group dynamics) and the Psychologist (who knows about emotional impact)."
  • The Twist: The Manager also whispers your specific preferences to the experts. They tell the Psychologist, "Remember, this user is extremely sensitive to dehumanizing language, so look for that specifically."

3. The "Expert Agents" (The Specialists)
The system uses different AI "experts" to analyze the content from different angles.

  • The Sociologist: Looks for discrimination or unfair treatment of groups.
  • The Linguist: Looks for rude words, insults, or aggressive tone.
  • The Psychologist: Looks for threats or things that might cause emotional distress.
  • The "Ghost" Agent: This is a special agent that acts exactly like you. If the experts disagree, the Ghost Agent steps in to say, "Based on what this user has told us before, they would find this harmful."

4. The "Synthesis Agent" (The Decision Maker)
Finally, a Synthesis Agent gathers all the opinions from the experts and the Ghost. It weighs their arguments against your personal profile and makes the final call: "Show this to the user" or "Hide this."

What Did They Find?

The researchers tested this system against the old "one-size-fits-all" methods and some of the smartest standard AI models currently used.

  • Better Accuracy: Their personalized team was 32% more accurate at predicting what a specific user would find harmful compared to the standard, generic systems.
  • Fewer Mistakes: The multi-agent team made fewer mistakes than a single AI trying to do everything alone. By breaking the problem down into smaller tasks (like having a specialist for insults and a specialist for threats), they avoided confusing the two.
  • Fast Learning: The system didn't need years to learn. Even after a user gave just a few pieces of feedback, the system started working much better than the generic version. It didn't take long to figure out your "taste."

The Big Picture: Who Decides?

The paper asks a big question: "Who decides what is harmful?"

Currently, the answer is "The Platform." They decide for everyone.
The PRISM system suggests the answer should be: "The User, with help from a smart team."

The authors argue that while platforms still need to keep the basics safe (like preventing violence), the day-to-day decisions about what is "too much" for a specific person should be up to that person. This system allows users to have a "filter" that fits their unique mental well-being, rather than forcing everyone to live under the same rigid rules.

In short: Instead of a security guard shouting "No one can say that!" to the whole crowd, PRISM gives every person a personal bodyguard who knows exactly what they can handle and what they can't, ensuring everyone feels safe in their own way.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →