Distinguishing True Spillover Trends from Surveillance Bias in Detected Cases
This paper argues that distinguishing genuine zoonotic spillover trends from surveillance bias requires robust infrastructure rather than just modeling, demonstrating through simulations and a rabies case study that while trend direction can often be recovered, magnitude estimates are highly unreliable and model disagreement should be treated as a critical signal for reporting.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
For decades, a quiet alarm has rung through the world of public health: the fear that diseases jumping from animals to humans are becoming more frequent. This idea has shaped research priorities and policy decisions, driving scientists to look for patterns in how often these "spillover" events occur. The logic is straightforward. If we see more cases being reported over time, it seems natural to assume the danger is growing. However, there is a hidden trap in this reasoning. The number of cases we know about depends entirely on how hard we are looking for them. If a country builds new clinics, trains more doctors, or simply starts paying closer attention to a specific disease, the number of reported cases will go up, even if the actual number of infections in the wild has stayed exactly the same. Distinguishing between a real biological surge and a statistical illusion caused by better detection is one of the most difficult challenges in modern epidemiology.
A team of researchers set out to test whether the tools scientists currently use can actually tell the difference between a real rise in disease and a rise in reporting. They focused on the methods used to analyze historical data, asking a simple but profound question: if the true number of animal-to-human infections remains steady, do our statistical models mistakenly tell us it is increasing? To find the answer, they did not look at real-world outbreaks first. Instead, they built a digital laboratory. They created thousands of simulated scenarios where they knew the exact truth: they programmed the computer to generate data where the number of spillover events was either rising, falling, or staying perfectly flat. Then, they layered on realistic imperfections, mimicking how detection effort changes over time, how it varies by distance from a testing lab, or how the same environmental factors that cause disease might also improve our ability to find it.
The researchers ran these simulations through three different types of statistical models, ranging from traditional regression techniques to modern machine learning algorithms. They wanted to see if these tools could recover the true trend hidden beneath the noise of imperfect data. The results were sobering. When the data was simple and the bias was straightforward, the models worked well. But as soon as the scenarios became more realistic—mimicking the complex, shifting nature of real-world surveillance—the models began to fail. Most notably, when the true number of infections was steady but the effort to find them changed, the models consistently lied. They frequently declared that the disease was on the rise when it was actually stable. This bias was particularly strong in scenarios where detection effort fluctuated, a common situation in the real world where resources wax and wane.
The study also revealed that even when different models agreed on the direction of a trend, they often disagreed wildly on the size of that trend. To prove this, the team applied their methods to a real dataset of bovine rabies cases in South Africa spanning twenty-seven years. All the models agreed that the number of cases was going down. However, when they tried to calculate exactly how fast it was dropping, the estimates varied by more than three times. One model suggested a slow decline, while another predicted a steep crash. This showed that while scientists might be able to agree on whether a trend is up or down, the precise numbers they produce are often unreliable and heavily dependent on the specific mathematical choices made by the researcher.
Perhaps the most critical finding concerns the reliability of the results we see in the news and scientific journals. The researchers found that when a model claims to have found an increasing trend, there is a significant chance it is wrong, especially if the underlying data comes from a system where detection effort is changing. In their simulations, when a model reported an increase, it was correct only about sixty percent of the time. This means that a large portion of the "rising spillover" claims made in the past could be artifacts of better detection rather than a genuine biological crisis. The study suggests that the field has been too quick to trust the magnitude of these trends. The authors argue that researchers should stop trying to pinpoint exact numbers of increase and instead focus on whether a trend is going up or down, while openly admitting the uncertainty.
Ultimately, the paper concludes that no amount of clever math can fully fix the problem if the data itself is incomplete. Statistical adjustments can only do so much when the surveillance infrastructure is weak or inconsistent. The researchers emphasize that the solution lies not in better software, but in better investment. To truly understand how zoonotic diseases are changing, we need sustained, active surveillance systems that search for cases regardless of whether they are reported by chance. Until we have data that is less biased by where and how we look, the most honest scientific approach is to treat claims of rising spillover with caution. The apparent increase in global spillover may be real, but it may also be a mirror reflecting our own growing efforts to find it, rather than a signal of a world becoming more dangerous.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.