Do established comorbidity scores predict in-hospital mortality equally well in women and men? A nationwide validation in 166.7 million German inpatient cases, 2010-2024
In a nationwide validation of 166.7 million German inpatient cases, three established comorbidity scores were found to be equally discriminative but systematically miscalibrated between women and men, supporting the need for sex-specific recalibration to ensure fair hospital benchmarking.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Hospitals constantly compare their performance, asking a fundamental question: are they keeping patients alive better than other hospitals? To answer this fairly, they must account for how sick a patient was before arriving. A patient who walks in with a broken leg and a healthy heart faces a different risk than one who arrives with a broken leg and a failing heart. To make this comparison, doctors use "comorbidity scores." These are simple numbers calculated from a patient's medical history, combining all their existing conditions into a single risk estimate. For decades, these scores have been applied to everyone using the exact same formula, regardless of whether the patient is a man or a woman. The assumption has been that the math works the same for both sexes. But in the world of medicine, where biology often differs between men and women, this assumption has rarely been tested on a massive scale. If the formula is slightly off for one group, it could skew the results of thousands of studies and distort the way hospitals are ranked, making some look worse or better than they truly are.
A team of researchers set out to test this assumption using a dataset of unprecedented size. They looked at nearly 167 million adult hospital admissions in Germany over a fifteen-year period, from 2010 to 2024. This was not a small sample; it was a complete record of almost every acute hospital stay in the country during that time. The researchers focused on three specific scoring systems that are used around the world to predict the risk of dying in the hospital. They wanted to see if these scores predicted death equally well for men and women. Specifically, they checked two things: first, whether the predicted risk matched the actual number of deaths (calibration), and second, whether the scores could successfully tell apart patients who would die from those who would survive (discrimination).
The results revealed a clear and consistent pattern. While the scores were generally good at ranking patients—meaning they could still distinguish between those at higher risk and those at lower risk—they did not predict the actual number of deaths equally for men and women. When the researchers looked at groups of patients with the same score, they found that men died more often than women. For example, in one of the scoring systems, men with a specific risk level were about 21 percent more likely to die in the hospital than women with the exact same score. This difference was not a fluke; it appeared consistently across all fifteen years of data and across almost every age group. The scores tended to underestimate the risk for men and overestimate it for women, or vice versa depending on the specific score, creating a systematic gap.
Interestingly, the ability of the scores to rank patients correctly remained very similar for both sexes. The difference in how well the scores separated survivors from non-survivors was so small it was almost negligible. The real issue was not that the scores failed to identify who was sicker, but that they miscalculated the actual probability of death for each sex. The researchers noted that for an individual patient, this difference is too small to change a doctor's immediate treatment decision. However, because these scores are used to compare entire hospitals and to guide national health policies, the error does not disappear. Instead, it accumulates. If a hospital treats more men than women, or vice versa, the standard formula might make that hospital look like it is performing worse or better than it actually is, simply because the math does not fit the specific mix of patients.
The study also explored why this might be happening. The data showed that men in the hospital population carried a higher burden of the most severe conditions linked to death, such as liver disease and heart rhythm problems, while women had higher rates of conditions like depression. The scores count these conditions, but they do not account for the fact that the same condition might carry a different weight or severity depending on the patient's sex. The researchers found that the pattern held true regardless of the year or the age of the patient, suggesting that the issue is built into how the scores were originally created and how they are currently used.
This massive analysis suggests that the long-standing practice of using a single formula for everyone is flawed when it comes to predicting hospital mortality. The authors conclude that these established scores are calibrated differently for men and women. While the scores still work well enough to rank patients, the systematic difference in predicting actual death rates means they introduce a bias into hospital comparisons and research. The study does not claim to have fixed the problem, but it provides strong evidence that the current tools need to be re-examined. To ensure fair comparisons between hospitals and accurate research, the authors suggest that these scores may need to be recalibrated specifically for men and women, or that sex should be added as a factor in the calculation. Until then, the numbers used to judge hospital performance may be quietly skewed by the sex of the patients they are meant to describe.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.