Evaluating missing-data workflows under hospital-structured missingness: a simulation-calibrated study in eICU and MIMIC-IV
This study demonstrates that hospital-structured missingness in multicenter electronic health records significantly biases mortality association estimates and undermines confidence interval coverage across conventional missing-data workflows, highlighting the critical need for cluster-level diagnostics, method-specific performance comparisons, and explicit missing-not-at-random sensitivity analyses.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern hospital, a patient's story is often told not just by what is written down, but by what is missing. When a doctor orders a blood test, the result appears in the electronic record. But if the test is never ordered, or if the machine fails to send the data, the record shows a blank space. For decades, researchers analyzing these massive digital archives have treated these blanks as simple errors to be fixed or ignored. They have assumed that a missing lab result is just a missing number, unrelated to the patient's condition or the hospital's habits. However, in reality, a missing value is often a message. It might mean a doctor decided a test wasn't necessary, or that a specific hospital lacks the equipment to run it. When these gaps are not random but follow patterns tied to the hospital or the patient's background, they can quietly distort the entire picture of what the data is saying.
This is the core challenge explored in a new study that looked at how researchers handle these missing pieces of information in critical care databases. The researchers wanted to know if the standard ways of filling in the blanks were actually changing the answers to important medical questions. They focused on a specific, high-stakes scenario: comparing the risk of death in the hospital for different racial and ethnic groups. The goal was not to prove that one group is biologically more fragile than another, but to see if the method used to clean the data could make it look like there was a difference where none existed, or hide a difference that was actually there. By testing these methods against a known truth created in a computer simulation, the study revealed that the choice of how to handle missing data is not just a technical detail; it is a decision that fundamentally changes who is being studied and what the results mean.
The researchers turned to two massive collections of real-world data from intensive care units in the United States. One database, known as eICU, gathered information from 206 different hospitals, capturing the diverse ways medical care is delivered across the country. The other, MIMIC-IV, came from a single, large academic hospital. They focused on eight common blood tests taken within the first day of a patient's stay, such as measurements of white blood cells, kidney function, and liver health. In the real world, these tests are not always available. In the single-hospital database, only about 24 percent of patients had all eight results recorded. In the multi-hospital database, the number was slightly higher at 36 percent, but the gaps were far from random. Some hospitals had data for almost every patient, while others had missing results for nearly everyone. In some cases, entire hospitals had no recorded values for specific tests like lactate or bilirubin, suggesting that the missing data was a structural feature of those specific medical centers, not a random glitch.
To see how this missingness affected the results, the team applied five different standard methods used by scientists to deal with incomplete records. One method simply threw away any patient with a single missing test. Another filled the gaps with the middle value found in the data and added a note to say the value was missing. A third method tried to weigh the remaining patients to represent the ones who were dropped. Two other methods used complex computer algorithms to guess the missing numbers based on the other information available for that patient, with one version trying to account for the specific hospital where the patient was treated. The researchers then used these five different versions of the data to calculate the difference in death rates between racial groups.
The results were startling. Depending on which method was used, the estimated difference in death rates between groups swung wildly. In one instance, when comparing Asian patients to White patients in the single-hospital database, the method that simply threw away incomplete records suggested Asian patients had a lower risk of death. Yet, a different method that filled in the missing numbers suggested they had a higher risk. The size of this swing was nearly three percentage points, a difference large enough to change the conclusion of the study entirely. In the multi-hospital database, the swings were just as dramatic, with estimates for Hispanic patients shifting from a lower risk to a higher risk depending on the workflow. The study showed that the method chosen did not just add a little noise to the data; it changed the direction of the finding.
To understand why this happened, the researchers built a computer simulation where they knew the exact truth. They created a virtual world of 60 hospitals and thousands of patients, programming the missing data to appear in specific patterns that mimicked the real world. Because they knew the true answer in this simulation, they could measure exactly how wrong each method was. They found that when missing data was tied to the hospital structure, even the most sophisticated methods struggled. The method that tried to account for the hospital context performed better than some, but it was not perfect. It still produced errors and confidence intervals that were too narrow to be trusted. The simulation also revealed that no method could fully fix the problem if the missing data was related to the patient's unrecorded severity in a way that the computer could not see. When the missingness was driven by factors the model could not know, every single method remained biased, producing estimates that were off by several percentage points.
The study also examined how the researchers' confidence in their results was affected. In science, a result is usually accompanied by a range of uncertainty, a margin of error that says how sure we are. The researchers found that some methods produced ranges that were far too tight, giving a false sense of precision. In the simulation, when the missing data was heavy and structured by hospital, the confidence intervals for some methods covered the true answer less than 30 percent of the time, even though they were supposed to cover it 95 percent of the time. This means that if a researcher used one of these flawed methods, they would be claiming to be very certain about a result that was actually quite shaky. The study highlighted that simply reporting a number without checking how the missing data was handled is like measuring a room with a ruler that shrinks and expands depending on the temperature.
One of the most significant findings was that the choice of method changed the population being studied. When researchers simply threw away patients with missing data, they were left with a group that looked different from the original crowd. In the real-world data, the patients who had all their lab results were often older or had different characteristics than those who did not. By filling in the gaps with averages or complex guesses, the researchers were effectively creating a new, synthetic population that might not exist in the real world. The study showed that the "standardized" risk difference they calculated was not a universal truth about the disease, but a specific answer to a specific question about a specific group of people defined by the method used to clean the data.
The researchers concluded that there is no single, magic solution for missing data in hospital records. The idea that a more complex mathematical model will automatically fix the problem was proven false in their simulations. Even a model that explicitly tried to account for the differences between hospitals did not guarantee an accurate result. The study argues that researchers must stop treating missing data as a technical cleanup step and start treating it as a fundamental part of the question they are asking. Before drawing any conclusions, they need to map out exactly where the data is missing and why. They must be honest about which group of people their results actually represent. And they must test their findings against different assumptions about why the data is missing, acknowledging that if the missingness is driven by hidden factors, the results could be wrong.
Ultimately, this work serves as a warning and a guide for the future of medical research using electronic records. It shows that the way we handle the blanks in our data can be just as important as the data we have. If we ignore the patterns of missingness, we risk building medical knowledge on a foundation that shifts with every change in methodology. The study does not offer a new, perfect tool to fix the problem, but it provides a clear map of where the pitfalls are. It suggests that the path forward requires transparency, rigorous testing, and a willingness to accept that some questions about health disparities may remain uncertain until we can better understand why the data is missing in the first place. The goal is not to find a perfect number, but to understand the limits of what our data can tell us.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.