← Latest papers
💻 computer science

The Course of News Events: A Comparison of Bottom-Up and Top-Down Approaches for Collecting Text-Based Data about Disasters

This paper compares top-down and bottom-up approaches for collecting text-based disaster data using German news articles on landslides, demonstrating that the choice of methodology significantly impacts event coverage and subsequent research outcomes in socio-environmental studies.

Original authors: Brielen Madureira, Andreas Niekler, Mariana Madruga de Brito

Published 2026-07-02
📖 4 min read☕ Coffee break read

Original authors: Brielen Madureira, Andreas Niekler, Mariana Madruga de Brito

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to create a complete history book of every landslide that happened around the world. You have two main ways to gather your stories, and this paper is like a detective comparing which method gives you the best, most accurate picture.

The Two Detectives: Top-Down vs. Bottom-Up

1. The Top-Down Detective (The "Checklist" Approach)
Think of this detective as someone holding a pre-written list of known landslides (called a "Disaster Inventory," like a famous global database called EM-DAT).

  • How they work: They look at their list, pick a specific landslide, and then go to the newspaper archives to find the story about that specific event.
  • The Catch: If a landslide happened but wasn't on their original list, this detective never looks for it. They only find stories about things they already know about.

2. The Bottom-Up Detective (The "Scavenger" Approach)
This detective doesn't have a list. Instead, they dive into the newspaper archives and use a computer program to scan thousands of articles. They look for patterns: "Hey, there are lots of articles about landslides in Country X happening around the same time. Let's group those together as one 'event'."

  • How they work: They build the list of events from the ground up, based entirely on what the news actually says.
  • The Catch: They might group things together that shouldn't be grouped, or they might find "events" that are just general discussions about landslides rather than a specific disaster happening right now.

The Experiment: German News about Global Landslides

The researchers tested these two methods using 55,000 German news articles about landslides from around the world. Here is what they found:

  • The Top-Down Detective missed a lot: When they tried to find news stories for the landslides on their official list, they only found a match for about 43% of them. This means for more than half the known landslides, there was no relevant German news story to be found near the time it happened.
  • The Bottom-Up Detective found a lot more (but with noise): This method found more than twice as many "events" as the official list had. However, only about 16% of these new "events" actually matched up with the official list.
  • The Overlap: Only about 760 events were found by both methods. This is the "safe zone" where everyone agrees a real event happened and was reported.

The "False Alarm" Problem

The paper points out a tricky issue with the Bottom-Up approach. Because the computer is just looking for words and dates, it sometimes creates "events" that aren't really disasters. The researchers found four types of "fake" events in the Bottom-Up pile:

  1. The "Real-Time" Event: A story about a landslide happening right now (Good!).
  2. The "History Lesson": A story talking about a landslide that happened 10 years ago, or a court case about an old disaster (Not a new event).
  3. The "General Discussion": An article about landslide risks, warnings, or scientific studies, with no specific disaster happening (Not a specific event).
  4. The "Metaphor": A story using the word "landslide" to describe a political defeat or a fictional movie scene (Completely wrong).

The Big Picture: Why It Matters

The researchers conclude that neither method is perfect on its own.

  • If you only use the Top-Down method, you miss many stories because you are limited by your original list. You might think a country had no news coverage, when in reality, the event just wasn't on your list.
  • If you only use the Bottom-Up method, you get a huge amount of data, but it's "noisy." You might count a history lesson or a metaphor as a real disaster, which skews your data.

The Takeaway:
To get the truest picture of how disasters are reported, you need to understand that news isn't just a mirror of reality; it's a filter.

  • The official lists (Top-Down) miss things that aren't considered "severe" enough to be recorded.
  • The news search (Bottom-Up) catches everything, including things that aren't actually new disasters.

The paper suggests that researchers need to be very careful about which "lens" they use. If you are studying how the media treats different countries, the choice of method changes the results. For example, the Bottom-Up method found more stories about the "Global South" (developing nations), while the Top-Down method found more stories about wealthy nations.

In short: Don't trust just one way of collecting data. You need to know the strengths and weaknesses of both the "Checklist" and the "Scavenger" to avoid drawing the wrong conclusions about the world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →