A Bayesian Prevalence Incidence Cure model for estimating survival using Electronic Health Records with incomplete baseline diagnoses
This paper proposes a Bayesian Prevalence Incidence Cure (PIC) model, a three-component mixture framework designed to accurately estimate survival, prevalence, and cure proportions from Electronic Health Records with missing baseline diagnoses, demonstrating its superiority over traditional Prevalence-Incidence models through simulations and a real-world application to Diabetic Macular Oedema.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery about a group of patients. You have a massive notebook of medical records (Electronic Health Records, or EHR), but the notebook is messy. Some pages are torn out, some entries are missing, and you don't always know exactly when a patient's condition started or if they will ever get better.
The paper by Pitt and Goudie introduces a new "detective tool" called the PIC Model (Prevalence-Incidence-Cure) to make sense of this messy data. Here is how it works, broken down into simple concepts.
The Three Types of Patients
In many medical studies, researchers usually assume everyone starts "healthy" and waits to see who gets sick. But in the real world, it's more complicated. The authors say you need to sort patients into three different buckets:
- The "Already Sick" (Prevalent): These patients already had the problem when they first walked into the clinic. Maybe they didn't have a test done that day, so the doctor didn't know they were sick until the next visit.
- The "Newly Sick" (Incident): These patients were healthy at the start, but later developed the problem.
- The "Unreachable" (Cured): These are the tricky ones. No matter how long you watch them, they will never get the problem (or never respond to the treatment). They are essentially "immune" or "non-responders."
The Problem with Old Tools
Before this paper, researchers used a tool called the PI Model (Prevalence-Incidence).
- The Flaw: The old tool assumed that everyone would eventually get sick if they waited long enough. It didn't have a bucket for the "Unreachable" (Cured) group.
- The Result: If you used the old tool on a group where some people would never get sick, the tool would get confused. It would try to force those "immune" people into the "Newly Sick" category, making it look like they would get sick much later than they actually would. This leads to wrong predictions about survival and treatment success.
The New Solution: The PIC Model
The authors built a 3-part mixture model (the PIC model) that acts like a smarter sorting machine.
- It handles the missing pages: It uses math to guess whether a patient with a missing initial test was "Already Sick" or "Newly Sick," rather than just guessing randomly.
- It handles the "Unreachable": It explicitly calculates the percentage of people who will never experience the event (the "Cure" proportion).
- It uses "Expert Intuition" (Bayesian Priors): The model allows researchers to feed in what they already know from past studies (like "we expect the median time to get better to be 10 months"). This acts like a compass to guide the math, especially when there isn't a lot of data.
The Detective Story: Diabetic Macular Oedema (DMO)
To test their new tool, the authors looked at a real-world dataset of 1,964 patients with a specific eye condition called Diabetic Macular Oedema (DMO). The goal was to see how long it took for patients to get their eyesight back to a "good" level.
- The Messy Data: Some patients had no eye test at the very start. Some had great vision immediately (they were "Already Sick" with good vision). Some never got better no matter what (the "Unreachable").
- The Comparison:
- The Old Tool (PI Model) said: "Everyone will eventually get good vision, but it might take a long time." It predicted that after 4 years, only 8% of people would still have bad vision.
- The New Tool (PIC Model) said: "Wait, about 18% of these people are 'Unreachable'—they will never get good vision with this treatment." It predicted that after 4 years, 18% would still have bad vision.
Why This Matters
The paper shows that the new PIC model is more accurate.
- It stops the "False Hope": The old model made it look like the treatment worked better than it did for everyone because it ignored the people who would never respond.
- It fixes the "Missing Data" confusion: It correctly identifies who was already doing well at the start, rather than counting them as people who "got better" later.
The Bottom Line
Think of the old model as a weather forecast that assumes it will rain eventually for everyone, even if some people live in a desert. The new PIC Model is a smarter forecast that realizes:
- Some people are already wet (Prevalent).
- Some people will get wet later (Incident).
- Some people live in a desert and will never get wet (Cured).
By acknowledging all three groups, the model gives doctors and patients a much clearer picture of what to expect, especially when the medical records are incomplete.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.