Bioinformatics sits at the exciting intersection where biology meets data science, using powerful computer tools to decode the vast complexity of living systems. From mapping the human genome to tracking how viruses evolve, this field transforms raw biological information into actionable insights that drive modern medicine and research forward without requiring a supercomputer to understand the basics.

On Gist.Science, we ensure you never miss a breakthrough by processing every new preprint in this category directly from bioRxiv. Our team provides both plain-language explanations and detailed technical summaries for each paper, making cutting-edge discoveries accessible to everyone regardless of their background.

Below are the latest bioinformatics papers added from bioRxiv, ready for you to explore with clarity and depth.

💻 bioinformatics

Ensemble-based genomic prediction for maize flowering-time improves prediction accuracy and reveals novel insights into trait genetic variation

This study demonstrates that an ensemble-based genomic prediction approach (EasiGP) significantly improves the accuracy of predicting maize flowering-time traits by leveraging the complementary strengths and diverse views of multiple individual models to offset prediction errors and better capture underlying genetic variation.

Tomura, S., Powell, O. M., Wilkinson, M. J., Cooper, M.2026-03-09
💻 bioinformatics

Benchmarking 80 binary phenotypes from the openSNP dataset using deep learning algorithms and polygenic risk score tools

This study benchmarks the performance of 29 machine learning algorithms, 80 deep learning models, and 3 polygenic risk score tools across 80 binary phenotypes from the openSNP dataset, revealing that machine learning approaches outperformed traditional tools for 44 phenotypes while polygenic risk scores were superior for the remaining 36.

Muneeb, M. -, Ascher, D., Myung, Y., Feng, S., Henschel, A.2026-03-09
💻 bioinformatics

anndataR improves interoperability between R and Python in single-cell transcriptomics

The paper introduces anndataR, an R package that enables seamless interoperability between R and Python in single-cell transcriptomics by allowing native reading and writing of H5AD files, conversion to and from SingleCellExperiment or Seurat objects, and rigorous testing to ensure long-term compatibility between the two languages.

Deconinck, L., Zappia, L., Cannoodt, R., Morgan, M., scverse core,, Virshup, I., Sang-aram, C., Bredikhin, D., Seurinck (…)2026-03-08
💻 bioinformatics

An Improved Dataset for Predicting Mammal Infecting Viruses from Genetic Sequence Information

This paper introduces a standardized, nearly doubled dataset of mammal-infecting viral pathogens with refined host labels to demonstrate that machine learning models achieve significantly better predictive performance for broader taxonomic ranks and when training and test sets share closer phylogenetic relationships, while highlighting the current limitations of generalizing these models to completely novel viral families.

Reddy, T., Schneider, A., Hall, A. R., Witmer, A., Hengartner, N.2026-03-08
💻 bioinformatics

The Stochastic System Identification Toolkit (SSIT) to model, fit, predict, and design experiments

The Stochastic System Identification Toolkit (SSIT) is a fast, flexible, open-source MATLAB package designed to model, simulate, and fit stochastic biochemical systems using diverse computational methods, enabling robust parameter inference, noise handling, and optimal experimental design for single-cell and other count-based data.

Popinga, A. N., Forman, J., Svetlov, D., Vo, H. D., Munsky, B. E.2026-03-08