Auditing missingness and input-outcome circularity in unsupervised clinical phenotyping: an empirical case study in hospitalized traumatic brain injury
This empirical case study on hospitalized traumatic brain injury demonstrates that unsupervised clinical phenotyping using routine scale data is highly sensitive to missingness handling and prone to input-outcome circularity, urging researchers to rigorously audit these methodological choices before assigning clinical meaning to derived subtypes.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to sort a messy pile of mystery boxes into neat groups based on what's inside. In the world of medicine, scientists often do this with patient data, a process called "unsupervised phenotyping." Instead of guessing which patients are alike, they use computers to find hidden patterns in routine check-up scores, like how well a patient remembers things, how they feel emotionally, or how easily they can move around. The goal is to discover natural "subgroups" of patients that doctors can use to give better, more personalized care.
However, there are two sneaky traps that can ruin the detective work. First, the data is rarely perfect; some boxes are missing labels or contents (this is called "missing data"). If you just guess what's missing, you might accidentally create groups that only exist because of your guesses, not because the patients are actually different. Second, there's a trap called "circularity." This happens if you use a specific clue to sort the boxes, and then act surprised when you find that the same clue is the main reason the boxes are different. It's like sorting a deck of cards by color, and then claiming you've discovered a magical new rule that "Red cards are red."
A team of researchers decided to test how well these detective methods hold up when applied to real-world hospital records of patients with traumatic brain injuries (TBI). They wanted to see if the groups of patients the computer found were real, stable, and useful, or if they were just a house of cards built on missing numbers and circular logic.
The Great Patient Sorting Experiment
The researchers took a look at 582 patients who had been hospitalized for brain injuries. They had five different "report cards" for each patient: tests for memory and thinking (MMSE and MoCA), tests for depression (GDS and Hamilton), and a test for how well they could do daily tasks like eating and dressing (Barthel ADL).
In a perfect world, every patient would have a score for every single test. But in the messy reality of a busy hospital, data goes missing. In this study, only 17.7% of patients (103 out of 582) had all five scores. The depression test and the daily living test were the most likely to be missing, appearing in nearly 40% and 30% of the records, respectively.
The team ran a standard computer program to sort these patients into groups. The program, using a method called "median imputation" (which basically fills in the blanks with the middle value of the existing data), decided that the best way to sort everyone was into five distinct groups. One of these groups looked particularly interesting: a "Hamilton-extreme" group of patients with very high depression scores.
But then, the researchers started their "stress test" to see if these groups were real or just an illusion.
The Stability Test: Do the groups hold together?
They tried to shake the data by randomly picking patients over and over again (a technique called bootstrapping). If the groups were real, the same patients should keep landing in the same groups every time. Instead, the groups fell apart like a wet sandcastle. The stability scores were very low (between 0.211 and 0.454), far below the 0.75 threshold needed to say, "Yes, this is a solid group." When they tried different sorting algorithms, the results barely agreed with each other.
The Missing Data Test: Does the method of filling blanks matter?
Next, they tried filling in the missing scores using a more complex method called "Multiple Imputation" (MICE), which creates 20 different versions of the dataset with different guesses for the missing numbers. When they ran the sorting on these 20 versions, the resulting groups were almost completely different from the original five groups. The agreement score was a tiny 0.157. This suggests that the original five groups weren't a discovery of the patients' true nature, but rather a reflection of how the researchers chose to fill in the blanks.
The "Complete Case" Test: What if we only look at perfect data?
They tried sorting only the 103 patients who had all their scores. When they did this, the computer didn't want five groups anymore; it preferred three. The "Hamilton-extreme" group that looked so important in the full dataset disappeared or changed completely. In fact, in this smaller group, the patients with high depression scores actually had better memory scores, which was the opposite of what the original analysis suggested. The researchers concluded that this specific "depressed" group was likely an artifact—a fake pattern created by the way the missing data was handled.
The Circularity Trap: Did we use the outcome to define the input?
Finally, they tested the "circularity" issue. In the original analysis, they used the "daily living" score (Barthel ADL) to help sort the patients, and then they checked if the groups were good at predicting the "daily living" score. Unsurprisingly, the groups were great at this! The statistical link was very strong.
But then, they removed the "daily living" score from the sorting process entirely and only used the thinking and mood tests. When they tried to predict the "daily living" score again using these new groups, the link vanished. The statistical score dropped from a strong 0.395 to a tiny 0.005. This proved that the "strong connection" they thought they found wasn't because the groups were naturally different in how they functioned; it was simply because they had used the function score to build the groups in the first place. It was a classic case of looking in the mirror and thinking you've discovered a new face.
The Takeaway
The study concludes that while unsupervised phenotyping sounds like a powerful tool to find new patient subtypes, it is incredibly fragile. In this specific case of traumatic brain injury, the "groups" the computer found were highly sensitive to how missing data was handled and were largely an illusion caused by using the outcome as an input.
The researchers suggest that before doctors start using these computer-generated groups to make clinical decisions, scientists must perform a rigorous audit. They need to check if the groups stay stable when data is resampled, if they survive different ways of filling in missing numbers, and most importantly, they must ensure they aren't accidentally using the result to define the group. Until these checks are done, the "subtypes" found in routine hospital data should be treated as interesting hypotheses to investigate, not as proven facts.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.