← Latest papers
💬 NLP

Not All Subjectivity Is the Same! Defining Desiderata for the Evaluation of Subjectivity in NLP

This position paper proposes seven user-centric evaluation desiderata for subjectivity-sensitive NLP models and identifies critical gaps in current research, such as the insufficient distinction between ambiguous and polyphonic inputs and the lack of interplay between different evaluation criteria.

Original authors: Urja Khurana, Michiel van der Meer, Enrico Liscio, Antske Fokkens, Pradeep K. Murukannaiah

Published 2026-03-31
📖 5 min read🧠 Deep dive

Original authors: Urja Khurana, Michiel van der Meer, Enrico Liscio, Antske Fokkens, Pradeep K. Murukannaiah

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a chef running a restaurant. For a long time, your job was simple: make a burger that tastes exactly like the "perfect" burger everyone agrees on. If you got it right, you got a gold star. This is how most Artificial Intelligence (AI) has worked so far. It tries to find the one "correct" answer, like solving a math problem or naming the capital of France.

But life isn't always a math problem. Sometimes, people disagree. Is a movie funny or offensive? Is a piece of art beautiful or disturbing? These questions don't have one right answer; they depend on who you are, your background, and your mood.

This paper, "Not All Subjectivity Is the Same!", argues that our current AI chefs are failing at these "opinion-based" dishes. They are trying to force a single flavor onto a plate that needs a whole buffet of tastes. The authors propose a new set of rules (called Desiderata) to teach AI how to handle human disagreement properly.

Here is the breakdown of their ideas, using some everyday analogies:

1. The Problem: The "Majority Vote" Trap

Currently, when AI learns from human opinions, it often takes a "majority vote." If 9 out of 10 people say a comment is "harmless," the AI learns to say "harmless."

  • The Issue: This silences the 1 person who felt hurt. In the real world, that 1 person might belong to a marginalized group. If the AI ignores them, it becomes a bully that only listens to the loudest voices.
  • The Goal: We need AI that can say, "Most people think this is fine, but some people find it hurtful because of their background," rather than just picking a winner.

2. The Two Types of "Confusion" (Ambiguity vs. Polyphony)

The authors make a crucial distinction between two types of tricky inputs. Think of this like a detective solving a case:

  • Ambiguity (The Missing Clue): The input is confusing because it's missing information.
    • Example: "I'm going to take this bitch out."
    • The Detective's Job: Is the speaker talking about a dog or insulting a person? The AI needs to realize, "I don't have enough context yet!" and ask for clarification.
  • Polyphony (The Chorus of Voices): The input is clear, but people genuinely disagree on the meaning based on their values.
    • Example: "I hate Boomers."
    • The Detective's Job: The sentence is clear. But for some, it's just a joke about age. For others, it's hate speech. There is no "missing clue"; there are just two valid realities existing at the same time. The AI needs to acknowledge both, not pick one.

The paper argues: Most AI treats these two things the same way. It tries to "solve" the polyphony (the disagreement) by picking a side, when it should actually be listening to the whole chorus.

3. The 7-Step Recipe for Better AI (The Desiderata)

The authors propose a 7-step checklist for building AI that handles opinions well. Imagine this as a training manual for a new AI chef:

  1. Recognize When: The AI must know when to stop looking for a single fact and start looking for opinions. (Is this a math problem or a debate?)
  2. Recognize Which: The AI must know what kind of opinion problem it is. (Is it a missing clue [Ambiguity] or a clash of values [Polyphony]?)
  3. Recognize Why: The AI should explain why people disagree. (Is it because of culture? Age? Personal experience?)
  4. Predict Accurately: It should guess the range of opinions correctly, not just the average.
  5. Calibrate: If the AI says, "There is a 50% chance this is offensive," it should be right 50% of the time. It shouldn't be overconfident.
  6. Represent All: It must give a voice to the minority. Don't just echo the majority; show the full spectrum of human thought.
  7. Express Without Annoying: This is the hardest part. If the AI lists 50 different opinions, the user gets bored. It needs to summarize the disagreement in a way that is helpful, not overwhelming. It's like a good moderator at a town hall meeting who summarizes the crowd's feelings without letting the meeting drag on for 10 hours.

4. What's Missing in Current Research?

The authors looked at 60 recent research papers to see how people are testing these AI models. They found some big gaps:

  • We are good at counting, bad at listening: Most papers just check if the AI got the "right" number of opinions (math). Very few check if the AI explained the disagreement well (communication).
  • We ignore the "Why": Researchers rarely ask why the AI made a specific prediction. Did it guess right for the right reasons?
  • We don't mix the types: We study "missing clues" and "clashing values" separately, but in real life, they often happen together.

The Big Takeaway

The paper concludes that for AI to be truly useful in society, it can't just be a robot that picks the "most popular" answer. It needs to be a diplomat.

A diplomat doesn't just tell you what the majority thinks; they understand why the minority feels differently, they know when to ask for more context, and they can explain the situation to you without making you feel confused or ignored.

The authors are calling for a new era of AI evaluation where we stop asking, "Did the AI get the right answer?" and start asking, "Did the AI understand the human experience behind the answer?"

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →