← Latest papers
📄 earth_science

Physically interpretable flood susceptibility mapping using spatial cross-validation and SHAP-derived environmental thresholds: A statistically validated framework

This study presents a statistically validated, physically interpretable framework for flood susceptibility mapping in Nepal's Triyuga River Watershed, utilizing spatial cross-validation and SHAP-derived environmental thresholds to demonstrate that a cumulative indicator of five critical hydro-geomorphic factors significantly enhances predictive accuracy and provides a transparent decision-support tool for risk management.

Original authors: Biplob Koirala

Published 2026-08-20
📖 6 min read🧠 Deep dive

Original authors: Biplob Koirala

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the foothills of the Himalayas, where the steep, rocky mountains meet the flat, fertile plains, water behaves in ways that can be both life-giving and devastating. This transition zone, known as the Siwalik-Terai, is a place where rivers, swollen by heavy monsoon rains, carry massive amounts of sediment from the eroding hills. As these rivers reach the flat land, they slow down, drop their load of sand and silt, and often spill over their banks, flooding the villages and farms that line their paths. For decades, scientists have tried to predict where these floods will happen next. They use computer programs that learn from past events, looking at factors like how steep the ground is, how close a location is to a river, and how much rain falls in an area. However, these computer models often act like black boxes: they give a result, but they do not explain why they reached that conclusion, and they sometimes make mistakes because they confuse patterns that happen to be near each other on a map with patterns that actually cause floods.

A new study focuses on the Triyuga River Watershed in southeastern Nepal, a region that perfectly illustrates these challenges. The researchers set out to build a flood prediction system that is not only accurate but also transparent and grounded in the physical reality of the landscape. Instead of relying on a single computer algorithm, they tested four different methods to see which one could best learn from the history of floods in the area. Crucially, they designed their testing process to ensure that the computer was not simply memorizing the map but was actually learning the rules of the terrain. By combining these rigorous tests with a technique that reveals exactly which environmental factors drive the predictions, they created a new way to identify dangerous areas. The result is a map that does not just show a probability of flooding, but highlights specific, measurable conditions—such as being within a certain distance of a river or sitting below a specific elevation—that must be present for a flood to occur.

The study began with a careful collection of data regarding where floods had actually happened between 2014 and 2025. The team gathered 240 confirmed locations of past flooding, verified through satellite images and historical records. To teach the computer what a safe area looks like, they had to create a set of "non-flood" points. This is a tricky step in flood science because true safe zones are hard to define; a place that didn't flood last year might flood next year. The researchers solved this by restricting their search for safe points to areas that were physically capable of flooding but simply hadn't this time. They looked at terrain that was low enough and close enough to the river to be at risk, but excluded the river channel itself. They then tested four different computer learning methods: a statistical approach called Logistic Regression, and three more complex machine learning models known as Random Forest, XGBoost, and Support Vector Machine.

To ensure the results were trustworthy, the researchers did not just split their data randomly. They used a method called spatial cross-validation, which divides the map into five distinct geographic blocks. The computer learns from four blocks and is tested on the fifth, rotating through all combinations. This prevents the model from learning the specific location of a flood rather than the underlying cause. When they ran these tests, the two tree-based models, Random Forest and XGBoost, performed the best, achieving an AUC of 0.905. Surprisingly, the simpler Logistic Regression model performed almost as well, suggesting that the rules governing floods in this area are not incredibly complex. The fourth model, Support Vector Machine, performed significantly worse, indicating it was not the right tool for this specific landscape.

The most significant breakthrough of the study came when the researchers asked the computer to explain its decisions. Using a technique called SHAP, which breaks down the model's logic, they discovered that five specific environmental factors were the primary drivers of flooding. The computer identified that being within 226 meters of a river, being below an elevation of 119 meters, having a slope flatter than 4.57 degrees, having a topographic wetness index above 8.33, and receiving more than 1,359 millimeters of annual rainfall were the critical thresholds. These are not vague guesses; they are precise boundaries where the risk of flooding changes dramatically. For instance, the distance of 226 meters corresponds closely to the typical width of the river's active floodplain, while the elevation of 119 meters marks the transition from the steep hills to the flat plains where water pools.

The researchers then combined these five rules into a single, easy-to-understand indicator. They created a map where every location is scored based on how many of these five critical conditions it meets. A spot that meets none of the conditions is considered very low risk, while a spot that meets all five is classified as an extreme hazard. The results were striking. When they checked this new map against the historical record of floods, they found that 85 percent of all past flood locations fell into areas that met at least three of these thresholds. Even more telling, locations that met all five conditions were 141 times more likely to have experienced a flood than areas that met none. This proved that flooding in this region is not caused by just one factor, like rain or distance to the river alone, but by the simultaneous presence of several specific conditions.

The study also produced a map showing where the computer was unsure. These areas of high uncertainty were found mostly along the edges where the flat floodplains meet the slightly higher terraces. In these transition zones, the environmental conditions are ambiguous, making it difficult for any model to be certain. The researchers suggest that these are the exact places where future monitoring efforts should be focused, as they represent the boundaries where flood risk is most dynamic. By identifying these zones, the study provides a clear guide for where to invest resources for better data collection and early warning systems.

This work offers a new template for managing flood risk in data-scarce mountainous regions. Unlike previous studies that relied on complex algorithms without explaining their logic, this approach provides a transparent set of rules that local planners can understand and use. The findings confirm that in the Triyuga River Watershed, the risk of flooding is governed by the physical shape of the land and the amount of rain, rather than by mysterious or unpredictable forces. The study demonstrates that by combining rigorous testing with explainable artificial intelligence, it is possible to create flood maps that are both scientifically robust and practically useful for protecting communities in the Himalayas. The framework developed here can be applied to other river basins in the region, offering a reliable way to identify dangerous areas and guide land-use decisions without needing extensive historical data or complex hydraulic models.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →