ROUGE : A database of disaster impacts in the Global South using Red Cross reports and Large Language Models
This paper introduces ROUGE, a new socio-economic disaster impact database for the Global South that leverages Large Language Models to extract detailed, non-monetary data from International Federation of Red Cross and Red Crescent Societies reports, thereby addressing existing gaps in geographic coverage and data bias.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world as a giant, bustling stage where nature occasionally throws a tantrum. Sometimes it's a storm that knocks over the set, a drought that dries up the props, or an earthquake that shakes the floorboards. For decades, scientists and aid workers have tried to keep a scorecard of these disasters to understand how much damage they cause and who gets hurt. But here's the problem: the scorecards we have so far are a bit like a library where only the books from wealthy countries are on the shelves. The stories from the Global South—the poorer, often more vulnerable regions—are missing, and the books we do have mostly count only money lost, ignoring the human stories, the broken homes, and the sick communities. To fix this, we need a way to read millions of messy, handwritten notes and official reports to find the hidden stories of pain and recovery. This is where a new kind of "super-reader" comes in: Large Language Models (LLMs). Think of these as incredibly smart AI detectives that can read thousands of pages of text in seconds, spotting patterns and pulling out specific facts that a human would take a lifetime to find. By teaching these AI detectives to read disaster reports, scientists hope to build a complete, fair picture of how natural hazards actually affect people around the world.
Enter ROUGE, a new database created by a team of researchers who decided to stop guessing and start reading the real stories. They took a massive pile of official reports from the International Federation of Red Cross and Red Crescent Societies (IFRC)—the global network of Red Cross volunteers that rushes to help after disasters—and fed them into an AI detective. These reports, written by field workers since 1919, are filled with details about what happened, who was hurt, and what was broken, but they are written in plain text, not neat spreadsheets. The researchers used a specific AI model (called meta-llama/llama-4-scout-17b-16e-instruct) to act as a digital archaeologist, digging through 2,610 reports from 2016 to 2025 to extract 15,147 specific impact records.
The result is a treasure map of disaster impacts that covers 776 unique events across 145 countries. Unlike older databases that mostly focus on money or only look at the country level, ROUGE zooms in. It captures 20 different types of impacts, ranging from the obvious (like "Human Deaths" or "Injured People") to the often overlooked (like "Homeless People," "Informal Settlements," or "Access to Food"). It even tracks how disasters affect things like schools, hospitals, and power grids. The team found that floods and storms are the most frequent troublemakers in their dataset, and that Africa and Asia are the regions where these impacts are most heavily recorded in the new database.
However, the researchers are careful not to claim this is a perfect, magic solution. They admit that AI can sometimes make mistakes, like guessing a date that wasn't there or mixing up two similar-sounding places. To make sure their work holds up, they tested the AI against a "gold standard" set of reports that humans had manually read and labeled. The results showed that the AI got about three-quarters of the facts right (a precision and recall score of roughly 0.73). It was particularly good at finding qualitative stories (like "people were displaced") but sometimes struggled with dense lists of numbers. They also compared their new data with existing databases like EM-DAT and IFRC GO and found that ROUGE often found numbers that the others missed—especially for things like "Missing People" or "Homeless People," where the new database provided unique data for up to 92% of the events it covered.
The paper suggests that while this new database is a powerful tool for filling in the gaps of our global disaster knowledge, it isn't a replacement for human judgment. The researchers built a system of "quality flags" to warn users when the AI might have been unsure or when a piece of data was missing a date or a location. They encourage anyone using the data to double-check the original text snippets if something looks suspicious. Ultimately, ROUGE suggests that by combining the tireless reading power of AI with the rich, on-the-ground stories of the Red Cross, we can finally start to see the full, unfiltered picture of how natural disasters impact the world's most vulnerable communities, moving beyond just counting dollars to counting lives and livelihoods.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.