← Latest papers
💻 computer science

Clinical explainable AI evaluations prioritize trust over effective human oversight: a scoping review

This scoping review of 49 clinical XAI studies reveals that evaluations overwhelmingly prioritize user trust and perceptions over critical oversight capacities like failure and bias detection, relying heavily on non-validated measures and rarely assessing performance on deployed systems.

Original authors: Nicolas Frey, Louis Agha-Mir-Salim, Linn Renner, Stefan Haufe, Lily Voge, Felix Balzer

Published 2026-09-24
📖 5 min read🧠 Deep dive

Original authors: Nicolas Frey, Louis Agha-Mir-Salim, Linn Renner, Stefan Haufe, Lily Voge, Felix Balzer

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In modern hospitals, doctors increasingly rely on computer programs to help diagnose illnesses and suggest treatments. These tools, built on complex mathematical models, can process vast amounts of patient data faster than any human. However, these models often work like a "black box," offering a recommendation without explaining how it reached that conclusion. To fix this opacity, researchers have developed a field called explainable artificial intelligence. The goal is simple: to provide the computer's reasoning alongside its answer, much like a doctor showing their work on a chalkboard. The hope is that by seeing the logic behind a suggestion, a clinician can decide whether to trust it, spot a mistake, or override it if the advice is dangerous. This ability to check and correct the machine is known as human oversight, and it is considered essential for safety in high-stakes medical decisions.

A new review of scientific research, however, suggests that the current way we test these explanations is missing the most critical part of the job. Nicolas Frey and his colleagues at Charité – Universitätsmedizin Berlin examined forty-nine studies published between 2015 and early 2026 that tested how doctors and medical trainees interact with these explainable tools. The researchers wanted to know if the studies actually proved that explanations help clinicians catch errors, or if they were merely measuring how much the clinicians liked the look of the explanation. They found a significant gap. While the studies frequently asked doctors if they found the explanations trustworthy or easy to understand, they almost never tested whether those explanations actually helped the doctors identify when the computer was wrong.

The review mapped out exactly what these forty-nine studies measured. The results showed a heavy focus on feelings and perceptions. In forty-two of the studies, researchers asked participants how much they trusted the system. In thirty-four studies, they asked if the explanations were useful, and in thirty-three, they asked if they were understandable. These questions rely on self-reporting, where a doctor simply says, "I feel confident in this answer." The researchers found that only a tiny fraction of the studies went a step further to test actual behavior. Just three studies tested whether a doctor could spot a specific failure in the computer's logic, and only one study tested if a doctor could recognize a systematic bias in the system's advice. Perhaps most strikingly, not a single study tested whether a doctor could correctly predict what the computer would do in a new situation, a skill that would be fundamental to true understanding.

The evidence suggests that the current research landscape prioritizes the feeling of safety over the mechanics of safety. Most of the studies took place in controlled, simulated environments rather than in real, busy hospitals. In these simulations, doctors were often shown correct and incorrect advice from the computer, but the studies rarely measured whether the doctors successfully rejected the wrong advice. Instead, they often measured whether the doctors agreed with the computer, without knowing if that agreement was with a right or a wrong answer. The researchers noted that a doctor might report high trust in a system and still blindly follow a dangerous recommendation, or they might feel confused by an explanation but still make the correct decision by ignoring the computer. Because the studies focused so heavily on the former—what the doctors said they felt—they provided little evidence that the explanations actually helped doctors catch mistakes when the computer was wrong.

Furthermore, the methods used to gather this data were often weak. More than half of the measurements relied on ad-hoc questions or unverified scales rather than established, scientifically tested tools. Only ten out of two hundred and fifty-three data points in the review used a validated instrument, which is a standard, reliable test for things like trust or satisfaction. The vast majority of the data came from simple surveys where doctors rated their experience on a scale. The review also highlighted that only two of the forty-nine studies looked at systems that were actually deployed and used in real clinical settings. The rest were experiments where the computer's behavior was controlled by the researchers, meaning we do not know if the explanations hold up when a doctor is under time pressure, tired, or dealing with a complex patient case.

The authors argue that this imbalance creates a false sense of security. They point out that trust is an attitude, while oversight is an action. A doctor can feel very trusting of a system but still fail to override a wrong recommendation. To truly know if an explanation is safe, researchers need to test specific behaviors: Can the doctor predict what the computer will say next? Can they spot a mistake when the computer is wrong? Can they recognize when the computer is biased against a certain group of patients? The review found that these specific, safety-critical skills were almost entirely absent from the current literature. The studies showed us how explanations feel to a clinician, but they did not show us if those explanations help a clinician do their job better when the machine fails.

The paper concludes that for explainable artificial intelligence to be truly useful and safe, the way we test it must change. Future studies need to move beyond asking doctors how they feel and start testing what they actually do. This means designing experiments where the computer's correctness is manipulated to see if doctors can distinguish right from wrong. It means using proven, reliable tools to measure attitudes rather than guessing based on simple surveys. Finally, it means testing these systems in the real world, over time, to see if the explanations continue to support good decision-making when the stakes are high and the pressure is real. Until these shifts happen, the medical community will have a clear picture of how much doctors like the explanations, but very little proof that those explanations actually keep patients safe.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →