← Latest papers
💬 NLP

Fault of Our Stars: Behavioral Drivers of Rating-Sentiment Incongruence

This study analyzes Sri Lankan tourism reviews to reveal that star ratings frequently diverge from textual sentiment due to specific behavioral patterns and contextual factors, demonstrating that ratings should not be automatically treated as ground-truth labels for sentiment analysis.

Original authors: Ramanaish Abaiyan, Ruththiragayan Sutharsan, Kusal Amantha, Anusan Krishnathas, Asma Rauff, Kovindarajah Sriyathurshan, Patalee Narasinghe, Nirasha Munasinghe, Nisansa de Silva, Sandareka Wickramanaya
Published 2026-06-25
📖 5 min read🧠 Deep dive

Original authors: Ramanaish Abaiyan, Ruththiragayan Sutharsan, Kusal Amantha, Anusan Krishnathas, Asma Rauff, Kovindarajah Sriyathurshan, Patalee Narasinghe, Nirasha Munasinghe, Nisansa de Silva, Sandareka Wickramanayake

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are at a restaurant. You leave a review on a popular app. You give the place 4 out of 5 stars, but in your written comment, you say, "The food was cold, the service was slow, and I almost got sick."

This is the exact problem this paper investigates. It asks: When people write online reviews, do their star ratings actually match what they say in their words?

The researchers from the University of Moratuwa in Sri Lanka decided to stop assuming the stars tell the whole story. Instead, they treated the stars and the words as two different people giving testimony, and they checked if those testimonies agreed.

Here is the breakdown of their findings, explained simply:

1. The "Two-Headed" Review

Think of every online review as having two heads:

  • Head A (The Stars): A quick, numerical score (1 to 5).
  • Head B (The Text): The actual story written by the person.

Usually, we assume Head A and Head B are best friends who always agree. This paper found that they are often strangers, or even enemies. In about 18.6% of the reviews they studied (roughly 1 in 5), the stars and the words told completely different stories.

2. The Six "Personas" of Mismatch

The researchers didn't just say "they disagreed." They found that these disagreements happen in six specific, predictable ways. Imagine these as six different types of confused reviewers:

  • The "Conservative Rater" (The most common): This person writes a glowing, happy story but gives a modest 3 or 4 stars. They are like a strict teacher who loves your essay but won't give you an 'A' because they want you to keep improving.
  • The "Obligatory 5-Star" (The second most common): This person writes a terrible review full of complaints but still gives 5 stars. They are like a polite neighbor who complains about your noisy dog but still waves and says, "Have a great day!"
  • The "Harsh Deflator": They write a positive story but give a low score.
  • The "Polite Inflator": They write a negative story but give a high score.
  • The "Frustrated Neutral": They write a negative story but give a neutral (3-star) score.
  • The "Positive Neutral": They write a positive story but give a neutral score.

The two biggest groups (Conservative and Obligatory) make up two-thirds of all the mismatches. This proves that the disagreement isn't random noise; it's a pattern.

3. Why Does This Happen? (The Drivers)

The researchers used computer models (specifically a type of AI called "Transformers," which are like super-smart reading machines) to figure out why people do this. They found four main reasons:

  • The "Museum Effect": Some places are just harder to rate than others. Museums had the highest rate of mismatch. Think of it this way: A museum visit is complex. You might love the art (positive text) but hate the crowded hallway and the expensive ticket (leading to a lower star rating). A beach or a park is simpler to rate, so the stars and words match up better there.
  • The "Expert" Reviewer: People who write lots of reviews (experts) are more likely to be "Conservative Raters." They seem to save their 5-star ratings for truly perfect experiences, even if they still write nice things about a "good" experience. Newbies are more likely to just give 5 stars for everything.
  • The Length of the Story: Longer reviews were more likely to have a mismatch. It seems that when people write a lot, they are expressing complex feelings that a simple 1-to-5 star scale can't capture.
  • Time Travel: Over the years (from 2010 to 2023), people have gotten slightly better at matching their stars to their words. The mismatch is slowly decreasing.

4. The Big Lesson for Computers (AI)

This is the most important part for the tech world.

For years, computer scientists have been training AI to understand human feelings (Sentiment Analysis) by using Star Ratings as the "Truth." They would feed the AI a review and say, "See? This is a 5-star review, so the AI must learn that this text means 'Happy'."

This paper says: Stop doing that.

If you train your AI using stars as the "truth," you are teaching it to learn from a liar. Because the stars often lie (or at least, they tell a different story than the words), the AI learns the wrong patterns.

The Analogy:
Imagine you are teaching a child to recognize a "happy" face.

  • The Old Way: You show them a photo of a person crying, but you tell the child, "This is happy," because the person is holding a 5-star trophy. The child learns that crying = happy.
  • The New Way: You look at the face itself (the text) to decide if they are happy, and you treat the trophy (the stars) as just a separate, sometimes unreliable opinion.

Summary

  • Stars and words often disagree.
  • Museums and expert reviewers are the biggest culprits.
  • The disagreement follows specific patterns, not random chaos.
  • Computer programs should not blindly trust star ratings as the "truth" when learning to understand human feelings.

The paper concludes that if we want to understand what people really think, we need to listen to their words, not just count their stars.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →