← Latest papers
💬 NLP

Approaches to Analysing Historical Newspapers Using LLMs

This study combines topic modeling, LLM-based sentiment analysis, and entity-graph visualization to examine how collective identities and political orientations were represented in early 20th-century Slovene newspapers, demonstrating the value of integrating scalable computational methods with critical discourse analysis for digital humanities research on noisy historical data.

Original authors: Filip Dobranić, Tina Munda, Oliver Pejić, Vojko Gorjanc, Uroš Šmajdek, David Bordon, Jakob Lenardič, Tjaša Konovšek, Kristina Pahor de Maiti Tekavčič, Ciril Bohak, Darja Fišer

Published 2026-03-27
📖 4 min read☕ Coffee break read

Original authors: Filip Dobranić, Tina Munda, Oliver Pejić, Vojko Gorjanc, Uroš Šmajdek, David Bordon, Jakob Lenardič, Tjaša Konovšek, Kristina Pahor de Maiti Tekavčič, Ciril Bohak, Darja Fišer

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a massive, dusty library filled with thousands of old Slovenian newspapers from the late 1800s and early 1900s. These papers are like time capsules, but they are also messy: the ink is faded, the paper is brittle, and the text was scanned by machines that sometimes misread letters (like turning an "m" into an "rn").

Reading every single article to understand how people felt about their neighbors, their country, or their enemies would take a human lifetime. That's where this study comes in. The researchers acted like digital detectives who used a team of super-smart AI assistants (called Large Language Models or LLMs) to read, sort, and analyze these newspapers for them.

Here is how they did it, broken down into simple concepts:

1. The Two Rival Newspapers (The "Team" vs. The "Opponent")

The researchers focused on two specific newspapers that were like two different sports teams with opposing fans:

  • Slovenec: The "Conservative-Catholic" team. They loved the Church and were very suspicious of liberal ideas and German influence.
  • Slovenski narod: The "Liberal-Progressive" team. They were big on nationalism, education, and were critical of the Church's power and German politics.

The goal was to see how these two "teams" talked about the same groups of people (like Germans, Austrians, or Slovenes themselves) and whether they were being nice, mean, or just factual.

2. The AI "Taste Testers" (Sentiment Analysis)

The researchers needed to know: When these papers mentioned "Germans," were they calling them friends or foes?

They tested four different AI models to see which one was the best "taste tester." They gave the AI a sentence and asked, "Is this positive, negative, or neutral?"

  • The Winner: One model, named GaMS3, was the best at the job.
  • The Quirk: However, the AI had a personality quirk. It was very good at spotting "neutral" facts (like "The train arrived"). But it was a bit shy about spotting "positive" feelings (it often missed compliments) and a bit too eager to spot "negative" feelings (it sometimes thought a neutral sentence was an insult).
  • The Lesson: The researchers learned that while the AI is a powerful tool, you have to know its blind spots. It's like a security guard who is great at spotting people who look suspicious but might miss a friendly wave.

3. The "Social Network" Map (Entity Graphs)

Once the AI read the papers, the researchers didn't just look at lists of words. They built a giant social network map.

  • The Nodes: Imagine dots on a map representing people (Germans, Slovenes), places (Vienna, Ljubljana), and topics (religion, politics).
  • The Lines: Lines connected the dots if they appeared together in the same article.
  • The Size: The bigger the dot, the more "emotional" the conversation was about that group.

What the map revealed:

  • Germans were often connected to big, angry dots (negative sentiment) in both newspapers.
  • Slovenes were often connected to mixed dots (some happy, some sad).
  • Regional groups (like "people from the coast") were often connected to tiny, faint dots, meaning they were just mentioned as facts without much emotion.

4. Zooming In: The "Close Reading"

The AI did the heavy lifting (the "distant reading"), scanning millions of words to find patterns. But the human researchers then zoomed in on specific parts of the map (the "close reading") to understand the story behind the data.

They found that the newspapers used specific tricks to build a sense of "Us vs. Them":

  • Othering: They often described Germans as the "enemy" or "outsider" to make Slovenes feel united.
  • Brotherhood: They described Croats as "brothers" to build alliances.
  • The "Us" Story: They emphasized that Slovenes should run their own schools and speak their own language in public.

The Big Picture Takeaway

This study is like using a satellite camera to see the whole forest (the AI analyzing millions of articles) and then sending a hiker into the woods to examine specific trees (the human researchers reading the actual text).

Why does this matter?
It shows that we can use modern AI to understand history, but we can't just trust the robot blindly. We need to combine the AI's speed with human wisdom to understand the messy, emotional reality of the past. The AI told them what was being said, but the humans explained why it mattered.

In short: They used a smart robot to read old newspapers, found out that the robot was good at spotting facts but bad at spotting compliments, and then used that data to draw a map showing how Slovenian people in the 1900s viewed their neighbors and their own identity.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →