A joint model for (un)bounded longitudinal markers, competing risks, and recurrent events using patient registry data
This paper proposes a novel Bayesian shared-parameter joint model that simultaneously handles multiple longitudinal markers (including bounded ones), recurrent events, and competing risks, demonstrating its superior performance through simulations and a practical application to cystic fibrosis patient registry data via the R package JMbayes2.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of medical research, doctors and scientists often track two different kinds of information to understand how a disease moves through a person's body. The first kind is a series of measurements taken over time, like checking a patient's blood pressure or lung capacity every few months. These are called longitudinal markers, and they tell a story of change, showing whether a patient is slowly getting better or worse. The second kind of information is about specific moments when something significant happens, such as a sudden hospital visit, a major complication, or even death. These are known as clinical events. For decades, researchers have tried to build mathematical models that link these two streams of data, asking how the slow, steady changes in a patient's health measurements might predict when a sudden, serious event will occur. This connection is vital because it allows doctors to anticipate problems before they become emergencies. However, real-life medical data is rarely simple. Patients often experience the same type of crisis multiple times, and they face different types of serious outcomes that compete with one another, such as the risk of dying versus the risk of needing a transplant. Furthermore, many health measurements are bounded, meaning they cannot go below zero or above a certain limit, which makes them difficult to analyze with standard tools that assume numbers can stretch infinitely in both directions.
A team of statisticians and epidemiologists has developed a new, more sophisticated way to connect these complex pieces of information. Their work focuses on cystic fibrosis, a severe genetic disorder that damages the lungs and digestive system. Patients with this condition are monitored closely for two key indicators: their body mass index, which reflects their nutritional status, and a measure of their lung function called the percentage of predicted forced expiratory volume in one second. This lung measurement is a percentage that typically ranges from zero to one hundred, though it can occasionally go slightly higher with effective treatment. Because this number is trapped within a specific range, standard statistical methods often struggle to model it accurately, sometimes producing impossible results like negative lung capacity. The researchers also needed to account for the fact that patients can suffer repeated lung infections, known as pulmonary exacerbations, and that they face two distinct, competing risks at the end of their journey: death from respiratory failure or receiving a lung transplant. Previous methods often forced researchers to ignore some of this information, such as treating multiple infections as a single event or combining death and transplant into one category, which obscured important details about the disease.
To solve these problems, the researchers created a unified statistical framework that can handle all these complexities at once. Instead of forcing the lung function data into a standard curve that allows for impossible values, they used a flexible mathematical shape designed specifically for numbers that stay within a set boundary. This approach ensures that the model only predicts realistic lung function values. They then linked this improved measurement model to the timing of repeated infections and the competing risks of death and transplant. By doing so, they could see how the trajectory of a patient's lung function and body weight influenced the likelihood of a future infection or a terminal event, while also accounting for the fact that a patient who receives a transplant is no longer at risk of dying from respiratory failure in the same way. The team tested this new method using a massive database containing health records from over twenty-three thousand individuals with cystic fibrosis in the United States, spanning nearly two decades. This dataset included millions of individual measurements and tracked thousands of hospitalizations, transplants, and deaths.
The results of applying this new model to the real-world data revealed a clearer picture of how the disease progresses. The analysis showed that a patient's risk of having a lung infection increases if they have had many infections in the past, confirming that the disease tends to cluster in episodes. More importantly, the model demonstrated that both lung function and body mass index are powerful predictors of future risks. Specifically, a faster decline in lung function was strongly associated with a higher risk of death, while the cumulative history of body mass index changes was linked to the risk of having another infection. The study found that the risk of death and the risk of needing a transplant are influenced differently by these markers, proving that treating them as a single outcome would have missed critical nuances. For instance, the model showed that a one-unit drop in lung function and a faster rate of decline each increased the hazard of death by distinct amounts, highlighting the importance of tracking the speed of change, not just the current value. The researchers also discovered that patients who were more frail, a term describing a general state of vulnerability, were significantly more likely to experience infections, receive a transplant, or die.
The study also served as a test to see if using the old, simpler methods would lead to wrong conclusions. When the researchers simulated data where the lung function was strictly bounded and then tried to analyze it with the traditional methods that assume numbers can go anywhere, the results were skewed. The older approach produced biased estimates, suggesting that the relationship between lung health and the risk of death was much weaker than it actually was. This confirmed that the new method is not just a minor tweak but a necessary improvement for accurate analysis. The team made their new model available to other scientists through a free software package, allowing researchers to apply this same rigorous approach to other diseases with similar complex data structures. By capturing the full story of a patient's health journey—the repeated setbacks, the competing dangers, and the bounded nature of their vital signs—this work provides a more precise tool for understanding disease progression. It offers a way to quantify the relationship between the most important clinical signs and life-threatening events with a level of detail that was previously impossible, potentially helping doctors manage these conditions more effectively in the future.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.