← Latest papers
💬 NLP

Automated stance detection in complex topics and small languages: the challenging case of immigration in polarizing news media

This paper demonstrates that both supervised language models and ChatGPT can effectively perform automated stance detection on the complex, lower-resource Estonian language regarding immigration, enabling the analysis of diachronic trends in polarizing news media.

Original authors: Mark Mets, Andres Karjus, Indrek Ibrus, Maximilian Schich

Published 2026-03-19
📖 5 min read🧠 Deep dive

Original authors: Mark Mets, Andres Karjus, Indrek Ibrus, Maximilian Schich

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to understand a massive, noisy town square where thousands of people are shouting about a single, controversial topic: Immigration. Some people are shouting "Welcome!" while others are screaming "Go away!" and many are just talking about the weather or the price of bread.

Your goal is to figure out who is saying what, and how the mood of the crowd changes over time. But there's a catch: you don't speak the language of the town (Estonian), and you don't have a team of 100 translators to read every single sentence.

This paper is about building a digital translator and mood-reader to solve this problem. Here is the story of how they did it, broken down into simple parts.

1. The Challenge: A Tiny Language and a Big Problem

Most computer programs that understand language are like giant libraries built for English. They have read millions of books. But for smaller languages like Estonian (spoken by only 1.1 million people), the library is very small.

The researchers wanted to see if they could teach a computer to understand the "stance" (the opinion) on immigration in Estonian news. This is tricky because:

  • The language is complex: Estonian words change their shape a lot (like adding different endings), making them hard for computers to recognize.
  • The topic is emotional: People use sarcasm, metaphors, and "dog whistles" (coded language) to hide their true feelings.
  • The data is scarce: There aren't enough pre-labeled examples to teach the computer the old-fashioned way.

2. The Experiment: Two Ways to Teach the Computer

The researchers tried two different methods to teach the computer how to sort the shouting crowd into three groups: Pro-Immigration, Anti-Immigration, and Neutral.

Method A: The "Apprentice" (Supervised Learning)

Think of this as hiring a human teacher to train a student.

  1. The Lesson: They took 8,000 sentences from news articles and had human students read them, labeling each one as "Pro," "Anti," or "Neutral."
  2. The Training: They fed these labeled sentences to a smart computer model (a type of AI called a Large Language Model). The model studied the patterns, like a student memorizing flashcards.
  3. The Result: The computer became a decent apprentice. It got about 66% of the answers right. It was really good at spotting "Anti" opinions but sometimes got confused between "Neutral" and "Pro."

Method B: The "Magic 8-Ball" (Zero-Shot ChatGPT)

This was the exciting new part. Instead of training a model with thousands of examples, they just asked a powerful, pre-existing AI (ChatGPT) a simple question in plain English:

"Here is a sentence about immigration. Is the writer for it, against it, or neutral? Just give me the label."

It's like asking a genius who has read the whole internet to guess the answer without any specific homework.

  • The Result: Surprisingly, the "Magic 8-Ball" performed almost exactly as well as the trained apprentice (65% accuracy).
  • Why it matters: This means you might not need to spend months and thousands of dollars labeling data for every new language. You can just ask the AI nicely, and it might do the job.

3. The Detective Work: What Did They Find?

Once they had their best "digital detective" (the trained model), they let it read 100,000+ sentences from two very different Estonian news sources over seven years (2015–2022):

  • Source A (Ekspress Grupp): A mainstream, general-interest news group (like a balanced newspaper).
  • Source B (Uued Uudised): A right-wing populist portal (like a political activist's blog).

The Findings:

  • The Split: As expected, the populist source was almost always "Anti-Immigration," while the mainstream source was mostly "Neutral."
  • The Shifts: The computer tracked how the mood changed during big world events:
    • 2015-2016 (Migration Crisis): Both sources talked more about it, but the populist source got louder and more negative.
    • 2018-2019 (UN Pact & Elections): The "Anti" stance spiked in the populist source right before elections, showing how politicians use these topics to rally voters.
    • 2022 (Ukraine War): This was the most interesting twist. When Russia invaded Ukraine, the mainstream source suddenly became much more "Pro-Immigration" (because Ukrainians were fleeing to Estonia). The populist source also shifted slightly, but they remained more skeptical.

4. The Limitations: Why It's Not Perfect

The researchers were honest about the flaws.

  • Context is King: Sometimes a single sentence is a trap. If a sentence says, "They criticize racism," the computer might think it's "Pro-Immigration." But if the sentence is actually quoting a villain saying that, the computer gets it wrong. It's like hearing a joke without knowing who is telling it.
  • Human Bias: The humans who labeled the training data had their own biases. If the teacher is biased, the student (the AI) will be too.
  • The "Black Box": With tools like ChatGPT, we don't always know how it got the answer, which makes it hard to trust 100%.

The Big Takeaway

This paper is a proof of concept. It shows that even for a small, difficult language and a messy, emotional topic, AI can be a useful tool for media monitoring.

It's like giving a researcher a pair of super-vision glasses. They can't read every single word in a newspaper, but with these glasses, they can quickly scan thousands of articles to see the big picture: Is the tone getting more angry? Is the topic shifting? Are the two sides of the political spectrum drifting further apart?

And the best news? You might not need a massive team of humans to build the glasses anymore; a simple prompt to a smart AI might be enough to get you started.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →