← Latest papers
📊 statistics

Functional forms in joint models for longitudinal and time-to-event data: A practical guide with application and interpretation

This paper provides a practical guide demonstrating that the choice of functional form in joint models for longitudinal and time-to-event data is a critical modeling decision that fundamentally shapes the interpretation of biomarker-risk associations, as illustrated through a comparative analysis of various association structures applied to glioblastoma trial data.

Original authors: Felix Boakye Oppong, Dimitris Rizopoulos, Thierry Gorlia, Nicole Erler

Published 2026-08-20
📖 5 min read🧠 Deep dive

Original authors: Felix Boakye Oppong, Dimitris Rizopoulos, Thierry Gorlia, Nicole Erler

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of modern medicine, doctors often track patients over years, watching how specific numbers in their blood or tissue change day by day. These numbers, known as biomarkers, tell a story about how a disease is behaving. At the same time, researchers are trying to understand when a patient might face a serious health event, such as a relapse or death. For a long time, scientists treated these two stories separately: one story about the changing numbers, and another about the timing of the event. But in reality, these stories are deeply intertwined. A patient's risk of a bad outcome is not just about where their numbers are right now, but how they got there. Did the numbers rise slowly over months, or did they spike suddenly last week? Did they bounce up and down wildly, or did they stay steady? Understanding the precise shape of this journey is crucial for predicting the future, yet figuring out exactly how to translate a complex medical history into a risk prediction has been a difficult puzzle.

A team of researchers set out to solve this puzzle by creating a practical guide for how to connect these two stories. They focused on a specific type of statistical tool called a joint model, which is designed to look at repeated measurements and survival times together. The core of their work was not about inventing new math, but about clarifying how to choose the right "lens" to view the data. They argued that the way a researcher decides to link a biomarker's history to a patient's risk is a critical scientific choice, not just a technical detail. If the wrong lens is chosen, the resulting conclusions about what drives the disease could be misleading. To demonstrate this, they applied their guide to real data from a major clinical trial involving patients with a type of brain cancer called glioblastoma. They used white blood cell counts, a standard measure of immune function, to show how different ways of looking at the same data could lead to different, and sometimes conflicting, understandings of the patient's risk.

The researchers identified several distinct ways to describe a patient's biomarker history, each capturing a different aspect of their medical journey. The most common approach simply looks at the current level of the biomarker at any given moment. It asks, "How high is the white blood cell count right now?" and assumes that this single number determines the risk. However, the team showed that this view misses the bigger picture. Two patients could have the exact same white blood cell count on a specific day, but one might have been rising steadily for weeks while the other was falling. If the speed of that change matters, the simple current-level view fails to distinguish between them.

To address this, the researchers explored more dynamic lenses. One approach looks at the speed of change, or the velocity, asking whether the biomarker is climbing or dropping at this exact moment. Another looks at the acceleration, or how quickly that speed is changing, capturing whether a rise is speeding up or slowing down. They also examined cumulative approaches, which look at the total exposure a patient has had to a biomarker over time. This is like calculating the total area under a curve, representing the burden of the disease over months or years, rather than just a snapshot. They also tested methods that focus on recent changes over a specific window, such as the last thirty days, to see if recent deterioration is more dangerous than long-term trends. Finally, they looked at variability, measuring how much a patient's numbers fluctuate around their average. A patient whose numbers swing wildly might be at higher risk than a patient with the same average but a steady line, because that instability could signal a body struggling to regulate itself.

When the team applied these different lenses to the data from the brain cancer trial, the results highlighted the importance of the chosen method, though the statistical evidence was often uncertain. When they looked only at the current white blood cell count, the link to survival suggested a possible positive association, but with considerable uncertainty as the confidence interval included zero. Similarly, when they looked at the cumulative exposure—the total burden of white blood cells over time—the association was statistically significant, yet the magnitude of the effect required careful interpretation. In other cases, looking at the rate of change or the variability provided different numerical estimates, but the wide confidence intervals for these measures indicated substantial uncertainty and weak evidence for a clear association in this specific dataset. The study found that these different functional forms are not interchangeable; they measure different biological realities. A large number in one model does not mean the same thing as a large number in another, because they are measuring different things, such as a current level versus a total history.

The researchers emphasized that there is no single "best" way to connect these data points. The right choice depends entirely on the biological question being asked. If a disease is driven by a sudden spike, looking at the current level might be best. If it is driven by years of exposure, a cumulative measure is necessary. If the instability of the system is the danger, then measuring variability is key. The paper concludes that researchers must carefully align their statistical choices with their scientific understanding of the disease. By treating the choice of how to link the data as a fundamental part of the scientific inquiry, rather than a hidden technical step, doctors and scientists can avoid drawing misleading conclusions. This clarity ensures that when a model predicts a patient's risk, it is based on the specific feature of their health history that actually matters, leading to more accurate predictions and better care.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →