← Latest papers
📊 statistics

Correcting heterogeneous diagnostic bias when developing clinical prediction models using causal hidden Markov models

This paper proposes a causal hidden Markov model framework to correct for heterogeneous diagnostic bias in clinical prediction models by estimating counterfactual diagnosis probabilities, thereby significantly improving calibration and reducing systematic errors in underdiagnosed groups as demonstrated in simulations and a chronic kidney disease case study.

Original authors: Jose Benitez-Aurioles, Ricardo Silva, Brian McMillan, Matthew Sperrin

Published 2026-05-08
📖 5 min read🧠 Deep dive

Original authors: Jose Benitez-Aurioles, Ricardo Silva, Brian McMillan, Matthew Sperrin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Spotlight" Effect

Imagine a doctor's office as a dark room, and patients are people walking around in it. Some people are wearing bright, reflective jackets (high-risk groups, like those with diabetes), while others are wearing dark, matte clothes (lower-risk or less-monitored groups).

In a perfect world, a doctor would shine a light on everyone equally to see if they are sick. But in the real world, the doctor tends to shine the bright spotlight only on the people in the reflective jackets because they look like they need help. The people in the dark clothes often stay in the shadows.

The result: The doctor finds many "ill" people in the bright light, but misses many "ill" people in the dark. If the doctor then tries to build a computer program to predict who will get sick in the future, the program learns a false lesson: "People in the dark clothes are healthy." It's not that they are healthy; it's just that no one looked at them closely enough to find the sickness.

This is what the paper calls heterogeneous diagnostic bias. It happens when certain groups (based on sex, ethnicity, or socioeconomic status) get tested less often, leading to "hidden" cases of disease.

The Solution: A "Time-Travel" Simulation

The authors propose a new way to fix this using a method they call a Causal Hidden Markov Model (HMM).

Think of a patient's health not as a single photo, but as a movie reel.

  • The Hidden Part: The movie has scenes where the patient is actually getting sicker (the "latent" disease stage), but the camera (the doctor) isn't rolling yet. We can't see these scenes in the real data.
  • The Visible Part: We only see the scenes where the camera is rolling (when a test is done).

The authors built a mathematical "time machine" (the HMM) that looks at the visible scenes and tries to guess what happened in the hidden scenes. It asks: "If this person in the dark clothes had been wearing a reflective jacket and been tested as often as the high-risk group, would we have found their sickness earlier?"

How the Method Works (Step-by-Step)

  1. Watching the Movie: The computer looks at the real-world data (who got tested, when, and what the result was).
  2. Filling in the Gaps: Using the HMM, it estimates the "hidden" movie. It guesses that many people who weren't tested were actually sick but undiagnosed, just like the people in the shadows.
  3. The "What If" Scenario: The computer then runs a simulation. It asks, "What if we forced the doctor to shine the light on everyone equally, just like they do for the high-risk group?"
  4. Rewriting the Script: Based on this simulation, the computer creates a new, "fair" version of the data. In this new version, the people who were previously missed are now marked as "diagnosed" because the simulation says they would have been found if the testing had been fair.
  5. Training the New Model: Finally, they train a new prediction model on this "fair" data. This new model doesn't learn the bias of the real world; it learns the true risk of the disease.

The Results: Fixing the Broken Compass

The authors tested this in two ways:

1. The Simulation (The Practice Run)
They created a fake world with 50,000 patients where they knew exactly who was sick and who wasn't.

  • The Old Way: A standard model looked at the data and thought the "underserved" group was much healthier than they really were. It was like a compass pointing North when it should be pointing East.
  • The New Way: Their HMM method corrected the compass. It realized the underserved group was just as sick as the others, just less tested. The "Observed vs. Expected" ratio (a measure of accuracy) went from a messy 1.34 (meaning the model was way off) down to 1.02 (almost perfect). Even when they broke some of their own rules (like making the tests imperfect), the new method still worked better than the old one.

2. The Real-World Test (Chronic Kidney Disease)
They applied this to real data from the UK (UK Biobank) regarding Chronic Kidney Disease (CKD).

  • The Discovery: They found that having Diabetes was the biggest factor in getting tested. People with diabetes were tested 10 times more often than those without.
  • The Bias: Because of this, standard models underestimated the risk for people without diabetes. They thought these people were safe, but the HMM showed they were actually at high risk, just missed by the system.
  • The Fix: When they used their new method to predict risk, the model stopped ignoring the non-diabetic group. The accuracy for this group jumped from a skewed 1.55 (way too many missed cases) to a perfect 1.01.

The Takeaway

The paper argues that we cannot just ignore "protected attributes" (like race or gender) to fix bias. Sometimes, those attributes tell us who is being tested and who is being missed.

By using this "time-travel" simulation, the authors created a model that separates true biological risk from systemic testing bias. It's like cleaning a dirty window: the view outside (the patient's true health) was always there, but the dirt (the testing bias) made it look blurry. Their method wipes the window clean, giving doctors a much clearer picture of who actually needs help.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →