← Latest papers
💬 NLP

Learning Moral Diversity: Modelling Individual Perspectives in Moral Classification of Texts

This paper proposes a method for modeling individual annotator perspectives in moral text classification by extending pretrained language models with annotator-specific features, demonstrating that accounting for subjective disagreement yields more accurate predictions and deeper insights than traditional approaches that aggregate labels into a single ground truth.

Original authors: Yi Ren, Lewis Mitchell, Matthew Roughan

Published 2026-06-23
📖 4 min read☕ Coffee break read

Original authors: Yi Ren, Lewis Mitchell, Matthew Roughan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to understand how people feel about a specific news story or a tweet. You ask 23 different friends to read it and tell you if it's "good," "bad," or "neutral" based on their personal moral compass.

In the past, researchers would take all 23 answers, mash them together, and say, "Okay, the majority says it's 'bad,' so that's the truth." They treated the disagreement between friends as just "noise" or mistakes to be ignored.

This paper argues that ignoring the disagreement is the real mistake.

Here is a simple breakdown of what the researchers did, using some everyday analogies:

1. The Problem: The "Average" Friend Doesn't Exist

The researchers looked at a huge collection of tweets where many people had labeled them with moral values (like "Fairness," "Loyalty," or "Purity"). They noticed that people often disagreed wildly. One person might see a tweet as a violation of "Authority," while another sees it as a champion of "Fairness."

Standard AI models tried to find a single "average" answer. The researchers say this is like trying to describe a complex painting by only looking at the average color of all the pixels. You lose the details, the texture, and the specific intent of the artist. By forcing a single "ground truth," the AI misses the rich, messy reality of how humans actually think.

2. The Solution: Giving the AI "Personalities"

The team built a new type of AI model. Think of a standard AI model as a generic translator that speaks one universal language.

Their new model is like a translator who has 23 different "personalities" or "masks."

  • They took a powerful, pre-trained AI (called BERT) that is already good at reading text.
  • They added a special "Annotator Layer" on top of it.
  • This layer acts like a personal filter for each of the 23 friends. It learns that "Friend #5" tends to be very strict about "Purity," while "Friend #12" is very sensitive to "Care."

Instead of asking, "What is the true moral value of this tweet?", the model asks, "What would Friend #5 say about this tweet?" and "What would Friend #12 say?"

3. The Results: Better Guesses and New Insights

When they tested this new model:

  • It got better at guessing individual opinions: The model became significantly more accurate at predicting what a specific person would label a tweet as, compared to the old "average" model. It improved accuracy by about 10%.
  • It revealed hidden patterns: Because the model learned the "personalities" of the annotators, the researchers could look at the data and see things like: "Oh, this group of friends always sees tweets as 'moral violations,' while that group is very cautious and rarely labels anything as moral."
  • The "Aggregated" Trap: When they tried to force their new, highly personalized model to give a single "average" answer (to mimic the old way), its performance actually dropped. This proved that the old way of averaging labels was hiding the fact that people simply see the world differently. The "average" answer wasn't necessarily the "best" answer; it was just a compromise that lost the nuance.

4. The Big Takeaway

The paper concludes that disagreement isn't a bug; it's a feature.

In the world of moral judgment, people don't just make mistakes; they have different perspectives. By building AI that respects and learns from these individual perspectives, we get a much clearer picture of human morality. The researchers suggest that future AI shouldn't just try to find the "one right answer," but should instead learn to understand the diverse ways people interpret the world.

In short: They taught an AI to stop pretending everyone thinks the same way, and instead, to learn how each specific person thinks. The result was a smarter, more honest model that understands the messy reality of human opinion.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →