← Latest papers
🧬 biology

Machine learning implementation for host-based infectious disease diagnostics

This systematic review of 21 studies (2020–2025) evaluates 54 machine learning models for host-based infectious disease diagnostics, finding that while Random Forest and neural networks are the most common algorithms, no single model is universally optimal, and future clinical translation requires standardized reporting, external validation, and context-aware algorithm selection.

Original authors: Samuel Harrington

Published 2026-09-03
📖 1 min read☕ Coffee break read

Original authors: Samuel Harrington

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Technical Summary: Machine Learning Implementation for Host-Based Infectious Disease Diagnostics

Problem Statement
The early identification of infectious disease etiology is critical for appropriate clinical treatment, yet uncertainty often leads to unnecessary antimicrobial exposure, contributing to the selection and spread of resistant organisms. Traditional diagnostic methods, such as bacterial cultures and targeted PCR, face limitations regarding speed, the requirement for a suspected organism, and the overlap of nonspecific inflammatory biomarkers across bacterial, viral, and noninfectious conditions. Host-based diagnostics, which evaluate the patient's biological response (e.g., transcriptomic signatures, cytokines, physiological signals) rather than the pathogen directly, offer a potential solution. However, the integration of these diverse data modalities with machine learning (ML) requires systematic evaluation to determine which algorithms are best suited for specific data types and how performance is currently reported in the literature.

Methodology
This study conducted a systematic review of literature published between January 2020 and August 2025. The search utilized four sources: Nature, PLOS, PubMed, and Elsevier, employing source-adapted search strings to capture studies applying ML to clinical patient data for pathogen detection, infection classification (e.g., bacterial vs. viral), or host-response diagnosis.

  • Selection Criteria: From 144 retrieved records, 51 passed an initial host-diagnostics screen. After excluding studies lacking suitable ML diagnostic models or focusing solely on in vitro systems, 21 articles met the final eligibility criteria.
  • Data Extraction: The review extracted 54 distinct model evaluations from the 21 articles. These evaluations were categorized by input modality (structured clinical variables, gene/molecular-expression measurements, medical imaging, physiological signals, and microscopy) and model class.
  • Model Focus: The analysis focused on four primary supervised model classes: Random Forest, Neural Networks, Support Vector Machines (SVM), and Logistic Regression.
  • Assessment: The author evaluated the consistency of performance reporting, noting that a formal risk-of-bias assessment (e.g., PROBAST) was not performed due to the substantial heterogeneity in study objectives and designs. Instead, methodological quality was evaluated descriptively during data synthesis.

Key Results

  • Model Distribution: Random Forest was the most frequently evaluated model class (23 of 54 evaluations, 42.6%), followed by Neural Networks (17, 31.5%), Support Vector Machines (9, 16.7%), and Logistic Regression (5, 9.3%).
  • Modality Alignment: Imaging studies generally favored Neural Network approaches, likely due to their ability to learn spatial features from high-dimensional inputs. Conversely, molecular-expression studies frequently utilized Random Forests.
  • Performance Reporting: Direct comparison of model performance across studies was not possible due to variations in input modalities, cohort composition, preprocessing, validation strategies, and reported endpoints.
    • Metrics: Accuracy and Area Under the Receiver Operating Characteristic Curve (AUC) were the dominant performance measures. 14 evaluations reported accuracy alone, 20 reported AUC alone, and 20 reported both.
    • Deficiencies: Only 16 evaluations (30%) reported additional clinically informative measures such as sensitivity, specificity, F1 score, or Matthews correlation coefficient.
  • Generalizability: The evidence did not support a universally optimal model class. Performance was heavily dependent on the alignment between the algorithm and the specific biological and technical properties of the input data.

Key Contributions

  1. Systematic Mapping: The paper provides a comprehensive mapping of the current landscape of host-based diagnostic ML, detailing the prevalence of specific algorithms across different data modalities.
  2. Reporting Analysis: It highlights a significant gap in reporting standards, specifically the over-reliance on Accuracy and AUC while under-reporting class-specific metrics (sensitivity, specificity) and calibration measures, which limits the clinical interpretability of these models.
  3. Contextual Model Selection: The review establishes that model selection should be driven by the structure and constraints of the input data (e.g., dimensionality, noise structure) rather than the pursuit of a single superior algorithm.
  4. Translational Framework: It outlines the necessary steps for clinical translation, including the need for external validation, transparent preprocessing descriptions, and the development of disease-agnostic outputs (e.g., bacterial vs. viral vs. noninfectious) rather than pathogen-specific classifiers.

Significance and Claims
The author claims that while host-based machine learning holds clear potential to support infectious disease classification and reduce unnecessary antimicrobial treatment, the field is currently hindered by methodological heterogeneity and inconsistent reporting. The paper asserts that no single model class is uniformly superior across all data types. Instead, the most immediate opportunity for advancement lies in improved standardization. The author argues that consistent reporting of preprocessing, validation, hyperparameter selection, and a broader suite of performance metrics (beyond just AUC) is essential to make studies reproducible and clinically interpretable. Future work must prioritize generalizable, disease-agnostic outputs and targeted biomarker panels that can be measured within clinically useful timeframes to transition these systems from retrospective demonstrations to valuable decision-support tools.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →