← Latest papers
📊 statistics

A fast and stable algorithm for non-parametric maximum likelihood estimation of survival functions for left-truncated and interval-censored data

This paper introduces a fast and stable product-limit style EM algorithm combined with a modified iterative convex minorant step to efficiently compute the non-parametric maximum likelihood estimator for survival functions using left-truncated and interval-censored data, demonstrating superior convergence and scalability over existing methods.

Original authors: Zachary Waller, Adele H. Marshall, Frank Kee, Felicity Lamrock

Published 2026-08-09
📖 4 min read☕ Coffee break read

Original authors: Zachary Waller, Adele H. Marshall, Frank Kee, Felicity Lamrock

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to figure out exactly when a specific event happens in a group of people, like when a secret club member finally decides to quit. But there's a catch: you don't get to see the moment they quit. You only get to peek at them at random times. Sometimes you look, and they are still there; the next time you look, they are gone. You know they quit sometime between those two peeks, but you don't know the exact second. This is called "interval censoring."

Now, add a second twist. Imagine you only start watching these people after they have already been in the club for a while. If someone quit before you started your watch, you never even knew they existed. This is called "left truncation." It's like trying to guess the lifespan of a tree, but you only start measuring it once it's already ten feet tall, and you only check on it every few years to see if it's still standing.

Scientists who study survival—like how long patients stay healthy or how long machines keep working—face this exact puzzle. They need a mathematical way to draw a map of time that shows the probability of an event happening, even when their data is full of these "I don't know exactly when" gaps and "I wasn't watching from the start" holes. The problem is that the old maps they used were incredibly slow to draw and sometimes got stuck in a loop, unable to find the best answer. If you want to know how confident you can be in these maps, you have to redraw them thousands of times, which makes the old, slow methods impossible to use for big, complicated real-world problems.

This paper introduces a new, super-fast detective tool called the "Product-Limit" (PL) algorithm, which works like a clever shortcut to solve this puzzle. The authors, researchers from Queen's University Belfast, realized that instead of treating the "missing time" as a messy mystery, they could reorganize the math to look more like a famous, simple method used for easier data. They call this a "reparameterization," which is just a fancy way of saying they changed the way they asked the question to make it easier to answer.

Think of the old way of solving this as trying to fill a giant, leaky bucket by pouring water in one drop at a time, hoping you eventually get it full. It works, but it takes forever, and if the bucket has a big hole (heavy truncation), the water might never stay in. The new PL algorithm is like realizing you can just plug the hole first and then pour the water in a steady stream. By treating the "start watching" times and the "stop watching" times as exact moments (which they are), and only using the complex math for the "I don't know exactly when" gaps, the new method skips the slow, repetitive steps.

The researchers tested this new tool against nine other existing methods using computer simulations. They created thousands of fake scenarios with different levels of missing data and "late starts." The results were clear: the new PL algorithm, especially when combined with a second step called "ICM," was dramatically faster and more stable than the others. In some tests, it was hundreds of times faster. While the old methods sometimes gave up or got stuck in a loop, the new method kept marching forward, finding the best possible map every time.

To prove it worked in the real world, the team applied their new algorithm to a famous dataset about older adults losing their ability to do daily tasks (like bathing or dressing). This data is tricky because the study only started watching people once they were already 65, and they only checked in every few years. The old methods took over 20 seconds to draw a map and sometimes got stuck after a million tries without finding the best answer. The new PL-ICM algorithm did the same job in a fraction of a second (0.003 seconds for women and 0.002 seconds for men) and found a more accurate map.

The paper suggests that this new approach is a game-changer for handling messy survival data. It doesn't just speed things up; it solves problems that other methods simply couldn't handle, allowing scientists to draw clearer, more reliable pictures of how time affects events, even when their data is full of gaps and late starts. The authors are confident that this method is ready to be used for complex studies, potentially helping researchers understand everything from disease progression to machine failure much more quickly and accurately.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →