← Latest papers
🤖 AI

Mod-Guide: An LLM-based Content Moderation Feedback System to Address Insensitive Speech toward Indigenous Ethnic and Religious Minority Communities

This paper introduces Mod-Guide, an LLM-based content moderation system enhanced by retrieval-augmented generation and a co-created corpus of insensitive speech from Bangladesh's Hindu and Chakma communities, which improves the detection of culturally harmful language by integrating minority perspectives and lived experiences into the moderation pipeline.

Original authors: Dipto Das, Achhiya Sultana, Ankit Singh Chauhan, Saadia Binte Alam, Mohammad Shidujaman, Shion Guha, Sunandan Chakraborty, Syed Ishtiaque Ahmed

Published 2026-06-12
📖 6 min read🧠 Deep dive

Original authors: Dipto Das, Achhiya Sultana, Ankit Singh Chauhan, Saadia Binte Alam, Mohammad Shidujaman, Shion Guha, Sunandan Chakraborty, Syed Ishtiaque Ahmed

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: A Translator for Cultural Feelings

Imagine the internet is a giant, noisy town square. In this square, there are two groups of people: the Majority (the most common voices) and the Minorities (smaller, often overlooked groups).

Sometimes, people in the Majority say things that aren't technically "hate speech" (like screaming slurs), but they are still hurtful. They might make jokes, assumptions, or comments that accidentally erase or disrespect the Minority's culture, religion, or history. It's like someone saying, "Why do you wear that weird hat? It's ugly," not realizing that hat is a sacred crown to the wearer.

The problem is that the "bouncers" (content moderation systems) at the town square are usually trained by the Majority. They often miss these subtle hurts because they don't understand the Minority's perspective. They might say, "That comment is fine," when the Minority feels deeply wounded.

This paper introduces Mod-Guide, a new tool designed to help the bouncers understand the Minority's point of view.

The Problem: The "Veil" and the "Missing Dictionary"

The authors use two powerful metaphors to explain why this happens:

  1. The Veil: Imagine a thick, invisible curtain separating the Majority and Minority. The Majority can't see the Minority's true feelings, and the Minority feels misunderstood.
  2. Hermeneutical Injustice: This is a fancy way of saying the Minority is missing a dictionary. They have experiences and feelings, but the "official language" of the internet (and the AI) doesn't have the words to describe why those experiences are painful.

Because of this, standard AI models (like the ones currently used by social media sites) often fail to spot "insensitive speech." They see the words, but they miss the meaning behind them.

The Solution: Building a "Community Library"

To fix this, the researchers didn't just ask an AI to guess. Instead, they went to the source.

Step 1: The Storytellers (The Corpus)
The researchers gathered 22 members from two specific minority groups in Bangladesh: the Hindu community (religious minority) and the Chakma community (Indigenous ethnic minority).

  • They asked these community members to share examples of online comments that hurt them.
  • Crucially, the community members didn't just say "this is bad." They wrote explanations. They explained why it was bad, using their own religious texts, history, and lived experiences.
  • Analogy: Imagine a library where, instead of just having a list of "bad words," you have a shelf of books written by the people who were hurt, explaining exactly why a specific sentence felt like a punch in the gut.

Step 2: The Tool (Mod-Guide)
The researchers built a tool called Mod-Guide. Think of it as a smart assistant for content moderators.

  • The Brain: It uses a powerful AI (GPT-4).
  • The Memory (RAG): This is the secret sauce. Instead of the AI guessing based on its general training, Mod-Guide forces the AI to open the Community Library (the corpus created in Step 1) before it answers.
  • The Role-Play: The AI is asked to wear different "hats" (personas) to give feedback. Sometimes it acts like a Teacher (trying to educate), sometimes a Mediator (trying to calm a fight), or a Judge (making a ruling).

How It Works in Real Life

Let's look at an example from the paper:

  • The Insensitive Comment: A user posts, "I don't believe in guarding idols; I brought my faith to guard them, not to protect statues." (This is offensive to Hindus who view idols as sacred).
  • Standard AI: Might say, "This is just an opinion, it's fine."
  • Mod-Guide (with the Library): The AI looks up the explanation from the Hindu community members. It sees that for them, idols are a bridge to God, not just statues.
  • The Result: Mod-Guide tells the user, "This comment might be hurtful because it dismisses the sacred nature of idols for many people. Here is how you could rephrase it to be more respectful."

What They Found

The researchers tested Mod-Guide against the standard AI and found:

  1. It Changes the Output: When the AI uses the "Community Library" (RAG), its answers are completely different from the standard AI. It becomes much more sensitive to cultural nuances.
  2. It's More Accurate: Experts from the minority communities reviewed the answers. They said the standard AI was often "shallow" or missed the point. The Mod-Guide answers were deeper and factually more accurate regarding their culture.
  3. Who Likes It?
    • Ethnicity Matters: People from the Indigenous (Chakma) minority found the feedback much more useful than people from the Majority.
    • Religion Didn't Matter as Much: Interestingly, whether someone was Hindu or Muslim didn't change how useful they found the tool; it was more about their ethnic background.
    • The "Teacher" and "Mediator" Hats: The feedback was most helpful when the AI acted like a teacher or a mediator, rather than a strict judge.

The Limitations (The Catch)

The authors are honest about what this tool can't do yet:

  • It's Small: The "Library" they built is small (132 examples). It's like having a dictionary with only a few pages. It works well for the specific groups they studied, but it might not cover every possible insult or every other minority group in the world.
  • It's Not Magic: The tool still needs human oversight. It can't replace human judgment entirely, especially for very complex or visual content (like images with text).
  • Specific to Bangladesh: The findings are specific to the Hindu and Chakma communities in Bangladesh. It doesn't automatically solve problems for minorities in other countries, though the method could be copied.

The Big Takeaway

This paper argues that to make the internet fair, we can't just rely on big, general AI models. We need to build local, community-specific libraries of knowledge.

By giving the AI a "dictionary" written by the people who are actually being hurt, we can help it understand the difference between a harmless comment and a deeply offensive one. It's about moving from punishment (banning people) to restorative justice (helping people understand each other and repair the relationship).

In short: Mod-Guide is a tool that teaches AI to listen to the people it's supposed to protect, rather than just guessing what they need.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →