← Latest papers
📄 health informatics

Evaluation of Deep Learning Based Early Warning Indicators of Epidemic Outbreaks

This paper evaluates deep learning-based bifurcation prediction models on real-world CDC influenza data across all 50 U.S. states, demonstrating that an expanding window configuration achieves high accuracy (90.09%) and recall (96.23%) in detecting epidemic tipping points, thereby validating their potential for real-time outbreak preparedness.

Original authors: Ayyorgun, B., Marathe, M., Adiga, A.

Published 2026-09-26
📖 1 min read☕ Coffee break read

Original authors: Ayyorgun, B., Marathe, M., Adiga, A.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Technical Summary: Evaluation of Deep Learning Based Early Warning Indicators of Epidemic Outbreaks

Problem Statement
In epidemic systems, detecting "critical transitions" or bifurcations—specifically the point where the basic reproduction number (R0R_0) crosses 1, shifting a disease from extinction to persistence—is vital for outbreak preparedness. While traditional statistical indicators (e.g., lag-1 autocorrelation, variance) can signal critical slowing down (CSD) near these transitions, they generally fail to predict future states or classify the specific type of bifurcation (e.g., fold, Hopf, transcritical).

Recent deep learning (DL) approaches, such as those by Bury et al. (trained on generic ODEs) and Chakraborty/Miry (trained on SIR-based stochastic simulations), have shown promise in detecting and classifying bifurcations. However, these models face significant limitations when applied to real-world data:

  1. Generalization Gap: Models trained on generic mathematical models or specific noise configurations often fail to detect bifurcations in real-world epidemic data characterized by varying noise levels and demographic factors.
  2. Lack of Rigorous Evaluation: There has been no systematic evaluation of these pre-trained DL models on real-world, state-level surveillance data spanning multiple geographic regions.
  3. Operational Uncertainty: It remains unclear how these models perform under different data windowing strategies required for online, real-time prediction.

Methodology
The authors developed a comprehensive pipeline to evaluate existing pre-trained DL bifurcation prediction models on a CDC-curated dataset of weekly influenza hospital admissions across all 50 U.S. states, Washington D.C., and national aggregates (July 2022 – early 2025).

  • Data Preparation:
    • Preprocessing: To match the training requirements of the Bury et al. models, the authors applied a strict preprocessing pipeline: (1) Detrending using a Lowess filter (span 0.2) to remove slow trends while preserving fluctuation structures; (2) Normalization by dividing residuals by the mean absolute value of the dataset.
    • Ground Truth Labeling: Since consistent serial interval data was unavailable for state-level RtR_t estimation, ground-truth bifurcation points were identified at the national level using the EpyEstim package (Bayesian renewal equation approach). For state-level evaluation, two manual bifurcation intervals were defined based on national trends: Interval A (t=108–118) and Interval B (t=127–140).
  • Models Evaluated:
    • Bury et al. (2021): A CNN-LSTM architecture trained on 500,000 time series from generic ODEs exhibiting fold, Hopf, and transcritical bifurcations.
    • Chakraborty/Miry (2024–2025): Models retrained on SIR-based stochastic simulations incorporating additive white, multiplicative environmental, and demographic noise.
  • Windowing Strategies: The study systematically evaluated how the models perform under different input tensor construction strategies for online prediction:
    • Rolling Window (Last 100): The original method using only the most recent 100 points.
    • Expanding Window: The window grows from the start of the series up to the current point (capped at 100).
    • Fixed-Length Moving Window: A sliding window of 100 points across the entire series.
    • Shorter Rolling Window: Length 50 with zero-padding.
    • Multi-Bifurcation Input: Feeding windows spanning multiple tipping points.

Key Results
The evaluation revealed that model performance is highly dependent on the windowing strategy and the specific bifurcation interval.

  • Performance Metrics: The Expanding Window strategy (retaining full historical context up to the current time, capped at 100 points) yielded the best performance across all metrics:
    • Accuracy: 90.09%
    • Precision: 72.86%
    • Recall: 96.23%
    • F1 Score: 82.93%
    • Specificity: 88.05%
  • Comparison of Strategies:
    • The original Rolling Window (last 100 points) achieved significantly lower accuracy (63.48%) and precision (17.35%).
    • Strategies that exposed the model to multiple bifurcation points within a single input tensor generally decreased accuracy for individual intervals but did not necessarily degrade aggregate performance if the first tipping point was sufficiently covered.
  • Interval Discrepancy: The models demonstrated high consistency in detecting the first major bifurcation (Interval A), with true positive rates of 51/53 locations. Performance dropped in Interval B (34/53 locations), likely because the timing of the second wave varied significantly by state, making it difficult to define a single evaluation window that encompassed all peaks without being too broad. The models performed best when the real-world dynamics closely resembled the training data distribution.

Significance and Claims
The paper claims that pre-trained deep learning models for bifurcation detection can generalize to real-world, state-level influenza data and are capable of online/real-time prediction, provided the input windowing strategy is optimized.

  • Contextual Importance: The study demonstrates that retaining the full temporal context of the series (via an expanding window) is superior to using only the most recent data points, suggesting that the "memory" of the system's approach to the tipping point is critical for the model's recognition capabilities.
  • Operational Viability: The findings indicate that these models are suitable for application in real-world surveillance, though the authors note that precision values in some configurations were driven by false positives, suggesting a need for threshold tuning in operational deployment.
  • Alternative Frameworks: The paper briefly discusses Markov Regime Switching (MRS) models as a promising alternative framework for detecting epidemic transitions, though a full comparative evaluation is reserved for future work.

The authors conclude that while these models show promise, their accuracy is contingent on the similarity between real-world bifurcation dynamics and the training data, and that careful selection of the input windowing strategy is essential for maximizing detection performance.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →