← Latest papers
🔬 physics

A Universal Convolution-Based Pre-processor to Correct the Prevalence-Incidence Gap in SIR, SEIR, and SIRS Modeling

This paper proposes a universal convolution-based pre-processor that corrects the fundamental methodological error of calibrating prevalence-based compartmental models (SIR, SEIR, SIRS) with raw incidence data, thereby eliminating systematic predictive biases and bridging the gap between clinical reporting and mechanistic epidemic forecasting.

Original authors: Jose de Jesus Bernal-Alvarado, David Delepine

Published 2026-02-02
📖 4 min read☕ Coffee break read

Original authors: Jose de Jesus Bernal-Alvarado, David Delepine

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: Mixing Up "New Arrivals" with "People Inside"

Imagine you are trying to predict how crowded a busy concert hall will get. You have two types of data:

  1. Incidence (New Arrivals): The number of people walking through the front door right now.
  2. Prevalence (The Crowd): The total number of people currently standing inside the hall.

The paper argues that for years, scientists have been making a massive mistake. They are taking the number of new people walking in (Incidence) and pretending it is the total number of people inside (Prevalence) to feed into their computer models.

The authors call this a "fundamental methodological error." It's like trying to guess how full a bathtub is by only looking at how fast the water is pouring in from the faucet, while ignoring that the drain is open and water is leaving at the same time.

Why This Matters: The "Peak" Mistake

When scientists use this wrong data, their models fail in two specific ways:

  • They get the timing wrong: The model thinks the crowd will peak (be most crowded) too early.
  • They get the size wrong: The model thinks the crowd will be much smaller than it actually is (underestimating the peak by up to 50%).

The paper shows that simply making the model more complicated—by adding more "rooms" to the building (like adding a waiting room for people who are sick but not contagious yet)—doesn't fix the problem. If you feed the wrong data (new arrivals) into a complex building, the whole structure collapses. The error just spreads to the new rooms.

The Solution: A "Universal Pre-Processor"

The authors propose a simple fix: a "pre-processor." Think of this as a special filter or a translator that you must use before you put any data into your model.

Instead of just counting the people walking in today, this filter looks at:

  1. How long people stay: If a sick person stays for 5 days, the filter remembers that the people who walked in 4 days ago are still inside.
  2. How fast they leave: It accounts for the fact that people recover and leave the "sick" room every day.
  3. Missing data: It acknowledges that not every sick person gets reported (some stay home). It adds a "correction factor" to guess the real number based on the reported numbers.

The Analogy of the Leaky Bucket:
Imagine the sick population is a bucket.

  • Incidence is the water pouring in.
  • Recovery is a hole in the bottom letting water out.
  • Prevalence is the actual water level in the bucket.

The old way of doing things was to say, "The water level is exactly the same as the water pouring in right now." This is wrong because the hole in the bottom changes the level.

The new method uses a mathematical recipe (a convolution) that looks at all the water that poured in over the last few days, subtracts the water that leaked out, and gives you the true water level.

The "Magic Formula"

The paper suggests a specific math formula to do this translation. In plain English, it says:

"To find out how many people are sick today, don't just look at today's new cases. Look at all the new cases from the past few days, but give less weight to the ones from a long time ago (because those people have likely already recovered)."

They also add a "reporting rate" to account for the fact that many sick people aren't getting tested or reported, ensuring the model doesn't underestimate the true size of the outbreak.

The Bottom Line

The paper concludes that no matter how fancy your epidemic model is (whether it's a simple SIR model or a complex one with immunity fading), you cannot skip this step.

If you want your predictions to be accurate, you must first translate "new daily cases" into "active sick people" using this specific filter. Without it, even the most advanced models will be looking at the wrong picture, leading to bad predictions about when the outbreak will peak and how big it will get.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →