Creation of the Estonian Subjectivity Dataset: Assessing the Degree of Subjectivity on a Scale
This paper introduces a new Estonian-language dataset for document-level subjectivity containing 1,000 texts rated on a continuous scale by human annotators, while also evaluating the feasibility and limitations of using GPT-5 for automated subjectivity scoring compared to human judgment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to measure how much "flavor" is in a bowl of soup. Some bowls are just plain water (completely objective facts), while others are loaded with spices, herbs, and the chef's personal opinion (completely subjective). For a long time, researchers in the world of computers and language (Natural Language Processing) have only been able to ask, "Is this soup spicy or not?" (Yes/No).
This paper is about building a new, much more precise measuring cup for the Estonian language. Instead of a simple "yes or no," the researchers created a dataset where they rated 1,000 different Estonian texts on a sliding scale from 0 to 100. A score of 0 is pure, unadulterated fact (like a weather report), and 100 is pure personal opinion (like a rant about a bad movie).
Here is the story of how they built it and what they found, explained simply:
1. The Recipe: Building the Dataset
The team wanted to see if humans could agree on these "flavor scores."
- The Ingredients: They gathered 1,000 texts. 300 were news and opinion pieces from the internet, and 700 were random web texts (like blogs, forums, recipes, and ads).
- The Tasters: They hired human "tasters" (annotators) to read these texts and give them a score between 0 and 100.
- The Pilot Test: Before tackling all 1,000, they tested the method on just 60 texts. It worked well! The tasters understood the scale and could tell the difference between a dry news report and a spicy opinion piece.
2. The Taste Test: Humans Don't Always Agree
When the team asked three different people to rate the same 1,000 texts, the results were a bit messy.
- The "Subjective" Problem: Just like people have different spice tolerances, people have different ideas of what counts as "subjective." One person might think a news story with a quote is neutral, while another thinks the quote makes it biased.
- The Correlation: The agreement between the humans was "moderate." It wasn't terrible, but it wasn't perfect. Sometimes, one person gave a text a score of 10 (very factual) and another gave it a 90 (very opinionated).
- The Re-Taste: To fix this, they picked the 250 texts where the humans disagreed the most and asked two of the tasters to rate them again. This time, the agreement got much better. It turned out that the order in which they read the texts mattered. If they just read a bunch of angry rants, they became hyper-sensitive to objectivity in the next text. It's like if you eat a super-sour lemon, the next piece of fruit tastes sweeter than it really is.
3. The Robot Taster: Enter GPT-5
The researchers then asked a question: "Can a super-smart computer (an AI called GPT-5) do this job as well as humans?"
- The Experiment: They let the AI rate all 1,000 texts.
- The Result: The AI was surprisingly consistent. If you asked it to rate the same batch of texts three times, it gave almost the same answers every time (unlike humans, whose moods and context changed their scores).
- The Difference: However, the AI and the humans didn't always see eye-to-eye.
- Quotes: When a news article quoted someone else, humans understood that the author was just reporting facts. The AI, however, saw the quote and thought, "Hey, this person is expressing an opinion!" so it gave the text a higher subjectivity score.
- Slang: Humans loved to give high scores to texts written in slang or forum speak because it felt personal and emotional. The AI was more focused on the facts inside the slang and gave it a lower, more "objective" score.
4. The Final Verdict
The paper concludes that:
- The Dataset is Real: They successfully created the first-ever Estonian dataset that measures subjectivity on a fine-grained scale (0–100) instead of just "yes/no."
- Humans are Flawed but Human: Humans can use the scale, but their scores are influenced by what they read right before and their own personal perspective.
- AI is a Helper, Not a Replacement: The AI (GPT-5) is great at being consistent and fast, but it "tastes" the text differently than humans. It misses the nuance of tone and context that humans catch.
In short: You can use the AI to quickly get a rough idea of how subjective a text is, but if you need to understand the human feeling behind the words, you still need a human to do the tasting. The AI and the human are like two different chefs using the same ingredients but following slightly different recipes.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.