← Latest papers
📄 health informatics

How Best to Explain Machine Learning Models to Clinicians: A User Study of Explanation Types

This user study involving 39 clinicians demonstrates that while attribution-based explanations most significantly enhance trust and understanding of machine learning predictions in clinical settings, nearly half of the participants prefer viewing multiple explanation types, suggesting that future implementations should prioritize attribution methods while offering diverse formats tailored to specific clinical roles.

Original authors: Brown, B., Oguss, M., Carey, K. A., Martin, J., Kotula, C. A., Nguyen, O. T., Akel, M., Wiegmann, D. A., Edelson, D. P., Mayampurath, A., Churpek, M. M., Craven, M.

Published 2026-07-10
📖 5 min read🧠 Deep dive

Original authors: Brown, B., Oguss, M., Carey, K. A., Martin, J., Kotula, C. A., Nguyen, O. T., Akel, M., Wiegmann, D. A., Edelson, D. P., Mayampurath, A., Churpek, M. M., Craven, M.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you've built a super-smart robot doctor that can look at a patient's data and predict if they are about to get sick. But here's the catch: the robot is a "black box." It gives you an answer, but it won't tell you why. It's like a magic 8-ball that says "Yes, danger!" but refuses to explain which crystal ball it used to see it.

To fix this, the researchers in this paper decided to test three different ways of explaining the robot's thinking to real-life doctors (nurses and physicians). They wanted to see which explanation style made the doctors trust the robot more, understand it better, and actually change their minds about what was wrong with the patient.

The Three "Flashlight" Styles

The team tested three different ways to shine a light on the robot's decision:

  1. The Attribution Flashlight (The "Scorecard"): This method gives every single piece of data a score. It's like a teacher grading a test and highlighting exactly which answers were right or wrong and how many points each one was worth. It says, "Your heart rate was high, and that added 4.5 points to the risk score."
  2. The Counterfactual Flashlight (The "What-If"): This one plays a game of "what if." It shows the doctor how to change the patient's numbers to get a different result. It's like a video game cheat code that says, "If you lower the heart rate and raise the blood pressure, the robot will stop screaming 'Danger!'"
  3. The Rule-Based Flashlight (The "Recipe"): This method gives a simple list of rules, like a cooking recipe. It says, "IF the lactate is high AND the breathing is fast, THEN the robot predicts danger."

The Big Reveal: The Scorecard Wins

After showing these three styles to 39 doctors and nurses, the results were clear. The Attribution Flashlight (the Scorecard) was the clear winner.

  • Trust: When doctors saw the scorecard, they trusted the robot's prediction more than when they saw the other two styles.
  • Understanding: The scorecard was the easiest to understand. Doctors felt they really "got it."
  • Influence: This is the most important part. The scorecard actually changed what the doctors thought was important. When they saw the scorecard, they were more likely to say, "Oh, I was wrong about what was causing the risk; the robot is pointing to something else."

The paper suggests that while the other two methods (the "What-If" and the "Recipe") were helpful, they didn't pack the same punch. The "What-If" style was actually the least liked in the free-text comments, with many doctors finding it confusing or less useful.

Doctors vs. Nurses: A Tale of Two Workflows

Here is a fun twist: the two types of medical professionals reacted differently.

  • The Physicians: Before seeing the explanations, they were already pretty interested in them. But after using the scorecard, their interest grew. It's like they thought, "I knew this was important, but now I see how it helps me make my big decisions."
  • The Nurses: They thought explanations were important from the start, but seeing them didn't really change their minds. The paper suggests this might be because nurses often follow strict, step-by-step protocols (like a checklist), so a complex explanation doesn't change their workflow as much as it does for a doctor who is making a diagnosis from scratch.

What the Doctors Wanted

The study also asked the doctors what they preferred to see on their screens.

  • More is Better (But Not Too Much): Nearly half of the doctors (46.2%) said they wanted to see all three types of explanations at once. They didn't want to choose just one; they wanted the scorecard, the "what-if," and the recipe all together to get the full picture.
  • Keep it Snappy: When it came to the details, the doctors wanted the "short and sweet" version. About 77% of them wanted to see only the top few most important features on the scorecard, rather than a long list of everything. Similarly, 84% wanted a short list of rules, not a novel-length recipe.

What the Paper Does NOT Say

It's important to know what this study didn't prove. The researchers did not say that these explanations make the robot doctor perfect or that they guarantee better patient outcomes in the real world. They only measured what the doctors said and thought during a survey. They also didn't test these methods on every possible type of robot doctor or every kind of hospital. The results are based on a specific robot trained on time-series data (like a patient's vitals over time) and a specific group of 39 doctors and nurses from two universities.

The Bottom Line

If you are building a robot doctor for a hospital, don't just give the human a single explanation. The paper suggests you should prioritize the Attribution Scorecard because it builds the most trust and understanding. However, since many doctors want to see everything, you should probably show them the Scorecard plus the other styles, but keep the lists short and easy to read. And remember, if you are talking to a nurse, the explanation might not shift their workflow as much as it does for a doctor, so tailor your approach accordingly.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →