A century of monthly rainfall in the Senegal River Basin (1895–2020): automated quality control, machine learning gap-filling, and the cost of network attrition
This study compiles a century-long, quality-controlled monthly rainfall archive for the Senegal River Basin and demonstrates that while machine learning effectively fills random data gaps, it is essential for correcting the significant wet bias introduced by terminal station attrition, which otherwise distorts the perceived magnitude of post-drought recovery.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Great Rain Detective Story
Imagine the Earth as a giant, breathing machine. One of its most vital organs is the weather, specifically the rain that falls on the ground. In science, this is called hydrology—the study of how water moves through the land, rivers, and clouds. To understand how this machine works, scientists need a long, unbroken diary of rainfall. They call these "time series." Think of it like a family photo album: if you have a picture every year for a hundred years, you can see how the family grew, who moved away, and what the weather was like during the big events. But what happens if someone tears out pages, writes the wrong numbers in the margins, or stops taking photos entirely? That is the problem scientists face in West Africa.
For decades, the region has suffered from a shrinking network of rain gauges (the devices that catch rain) and old records filled with typos. This matters because the Senegal River Basin is a lifeline for millions of people. It provides water for drinking, farming, and keeping huge reservoirs full. If the "diary" of the rain is broken or full of lies, the people managing the water might make terrible mistakes—like opening a dam when it's about to flood, or closing it when the crops are dying. The big question is: Can we fix the broken diary? Can we use modern computer tricks to fill in the missing pages and correct the typos without inventing new rain that never fell?
The Century-Long Rain Rescue Mission
This paper is a massive rescue mission for the Senegal River Basin, a vast area shared by Guinea, Mali, Mauritania, and Senegal. The authors, a team of researchers, decided to gather every single monthly rainfall record they could find from 1895 to 2020. They didn't just collect them; they built a super-smart, automated system to clean up the mess. Imagine a librarian who not only finds books but also spots when someone has written "1000" instead of "10.00" in a ledger, or when two different people wrote two different stories for the same day.
The team assembled a library of 138 stations and over 80,000 observations. But before they could use this data, they had to play "spot the error." Their automated system found that 1.4% of the records were messed up. Some errors were wild: one station in Mali recorded 111,622.8 mm of rain in a single month (that's more than 111 meters of rain, which is physically impossible!), and another recorded 90,120 mm. The system realized these were just decimal point mistakes and fixed them by dividing by 100. The lesson here is scary but important: leaving just one of these errors uncorrected can shift the entire region's rainfall history by tens of standard deviations. It's like if one person in a room of 100 suddenly weighed 10 tons; the average weight of the whole room would look completely wrong.
Once the data was clean, the researchers asked the next big question: How do we fill in the gaps? Over the years, many rain stations closed down, leaving huge holes in the record, especially after the 1990s. The authors tested six different methods to guess what the rain was during those missing months. They used everything from simple math (looking at what your neighbors got) to fancy machine learning (computer programs that learn patterns).
Here is where the story gets interesting. The researchers ran thousands of simulations, pretending to hide parts of the data and then seeing which method could guess it best.
- For random gaps (like a few missing months here and there) or block gaps (missing whole years), the simplest methods won. A classic technique called the "normal-ratio" method and a simple linear regression were the champions. They recovered the missing rain with a skill score (Kling-Gupta efficiency) near 0.90. The fancy machine learning models, like Random Forest and XGBoost, did just as well as the simple ones but never beat them. It turns out, when you have a dense network of neighbors, you don't need a super-computer; you just need a good map of who lives next door.
- However, there is a catch. The paper found that when the gaps happen at the end of the record (because stations closed down during a time when the climate was changing), the simple methods fail. They start to guess that it was wetter than it actually was. This is called a "wet bias." Because the models were trained on wetter years from the past, they kept guessing "it's raining!" even when the drought had started. In these specific "terminal gaps," the bias could reach 9.4 mm per month with simple methods.
But the machine learning models saved the day in this specific scenario. When the researchers pooled all the stations together and gave the computer a "calendar year" as a hint (telling it, "Hey, we are in the 1990s, things are getting drier"), the machine learning models could fix the bias. They cut the error down to just 2.0 mm per month. So, machine learning isn't magic for every problem, but it is the only tool that works when the climate is shifting and the data network is falling apart.
The team also checked their work against satellite data. The satellites confirmed that the region is indeed recovering from a long drought that lasted from 1968 to 1993. However, the satellite data showed the recovery was gentler than the rain gauge data suggested. This proved that the shrinking network of gauges was making the recent rain look stronger than it really was—a classic case of "attrition bias."
In the end, this paper gives us a clean, corrected, century-long diary of rain for the Senegal River Basin. It tells us that for routine missing data, simple, neighbor-based math is still the king. But when the network shrinks and the climate changes, we need smarter, pooled machine learning models to avoid fooling ourselves. The authors have released this cleaned-up database to the public, giving the four countries sharing the river a trustworthy foundation to manage their water, plan for floods, and negotiate peace. It's a reminder that in science, cleaning up the past is just as important as predicting the future.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.