← Latest papers
💻 computer science

On the Disagreement in Perturbation-based xAI -- Benchmarking Perturbation Choices for Flood Detection from SAR Images

This paper investigates how critical parameter choices in perturbation-based explainable AI, specifically patch geometry and perturbation type, significantly alter relevance maps for flood detection in SAR imagery, highlighting the necessity of rigorously evaluating these settings to ensure robust and faithful model interpretations.

Original authors: Anastasia Schlegel, Ronny Hänsch

Published 2026-07-17
📖 4 min read☕ Coffee break read

Original authors: Anastasia Schlegel, Ronny Hänsch

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot that can look at a picture and tell you exactly where a flood is happening, even if the photo was taken from space through thick clouds. This robot is a "deep learning model," a type of computer brain that has learned to recognize patterns by looking at thousands of examples. But here's the catch: the robot is a bit of a black box. It gives you the right answer, but it doesn't tell you why. Did it see the dark, wet water? Or did it just get confused by a shadow on a mountain that looks dark?

To solve this mystery, scientists use something called "Explainable AI" (xAI). Think of it like a detective trying to figure out what clues the robot used. One popular method is "perturbation." Imagine you have a photo of a flood, and you start covering up tiny pieces of it with a gray square, one by one, to see if the robot stops recognizing the flood. If covering a specific spot makes the robot say, "Wait, I don't think this is a flood anymore!", then that spot must have been important. It's like playing a game of "What if?" to see what the robot cares about. But what if the game itself is rigged? What if the size of your gray square or the color you use to cover the spot changes the answer? That is the big question this paper asks.

The researchers behind this study decided to put this "What if?" game to the test, specifically using radar images of floods. They wanted to see if changing the rules of the game—like using a tiny square versus a giant one, or covering a spot with a gray color versus a random pixel from another photo—would change the robot's "reasoning." They found that the answer is a loud, resounding "Yes."

In their experiments, the team discovered that the robot's "explanation" is incredibly sensitive to how you play the game. They tested different shapes and sizes for the patches they used to cover the image, ranging from small 8x8 pixel squares to large 128x128 pixel blocks. They also tried different ways to fill in the covered spots: some were filled with the darkest possible color, some with the average background color, and some with random pixels from other images.

The results were surprising and a bit chaotic. When they used a large, blocky square to cover the image, the robot's "heat map" (the visual guide showing what's important) became blurry and coarse, like a low-resolution video. It couldn't tell the difference between a thin river and a wide lake. But when they used smaller patches or shapes that matched the natural curves of the land (like superpixels), the robot's reasoning became much sharper and more detailed.

Even more confusing, the type of color they used to cover the spot changed the story entirely. For example, when they covered a spot with the "minimum" (darkest) value, the robot seemed to think the background was the most important thing, ignoring the water. But when they used the "average" background color, the robot suddenly started focusing on the water, and in some cases, it even flipped its opinion, saying the water was helpful in one scene but distracting in another. It's as if the robot was saying, "Oh, I only care about the water if you cover the ground with this specific color!"

The researchers also checked if these explanations were actually telling the truth. They used a "deletion test," where they removed the parts of the image the robot said were most important to see if the robot actually got confused. They found that for some methods, the robot didn't get confused until they removed almost the entire image, meaning the explanation was misleading. For other methods, the robot lost its ability to detect the flood immediately, suggesting that explanation was more honest.

The main takeaway is that there is no single "correct" way to ask the robot what it's thinking. The choice of how you perturb the image—how big your "cover-up" is and what you put in its place—completely steers the explanation. The authors suggest that we can't just pick a method at random and trust the result. Instead, we have to be very careful, like a scientist choosing the right tool for a job. If we want to understand a flood detection model, we need to realize that the explanation we get is a mix of the model's brain and the specific rules we used to probe it. Without being careful about those rules, we might end up trusting a story that the robot didn't actually tell.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →