Bioinformatics sits at the exciting intersection where biology meets data science, using powerful computer tools to decode the vast complexity of living systems. From mapping the human genome to tracking how viruses evolve, this field transforms raw biological information into actionable insights that drive modern medicine and research forward without requiring a supercomputer to understand the basics.

On Gist.Science, we ensure you never miss a breakthrough by processing every new preprint in this category directly from bioRxiv. Our team provides both plain-language explanations and detailed technical summaries for each paper, making cutting-edge discoveries accessible to everyone regardless of their background.

Below are the latest bioinformatics papers added from bioRxiv, ready for you to explore with clarity and depth.

💻 bioinformatics

ML4SD: Leveraging Machine Learning and High-Throughput Search Algorithms for an Iterative Growth-Coupled Design Innovation

The paper introduces ML4SD, an active-learning Design-Build-Test-Learn cycle that integrates machine learning with a novel high-throughput algorithm (gcSwarms) to efficiently identify growth-coupled microbial strain designs, achieving significantly higher carbon yields with far fewer experimental iterations than traditional search methods.

Gargantilla Becerra, A., Nogales Enrique, J.2026-09-16
💻 bioinformatics

A Turing-Style Test for In-Silico Antibodies: How Sampling Mode Makes WGAN-GP Beat VAE in the Wet Lab

This study demonstrates that while Wasserstein GANs outperformed Variational Autoencoders in generating experimentally viable de novo antibodies (99% vs. 5% success), this disparity was primarily driven by the sampling strategy—unconditional generation versus latent seeding with known antibodies—rather than inherent architectural superiority.

Kummer, A., Mahmoudinobar, F., Liu, W., Davis, J. W., Ma, E. J., Kumar, S.2026-09-16
💻 bioinformatics

Uncertainty-Aware Model Selection with a Calibrated Probability-Generating-Function-Based Bayesian Information Criterion

This paper introduces an uncertainty-aware model selection rule for stochastic gene-expression models that enhances the conventional PGF-BIC by incorporating a data-driven threshold, derived via influence functions and Cantelli's inequality, to account for sampling uncertainty and prevent the over-selection of complex models without sacrificing computational efficiency.

Wang, Y., Shu, Z., Gao, F., Cao, Z.2026-09-16
💻 bioinformatics

Interpretable Machine Learning Reveals Complementary Age-Related Signatures in the Oral and Gut Microbiome

By employing SHAP-based interpretability on paired oral and gut microbiome data, this study reveals that while combined models do not surpass the perfect predictive accuracy of gut data alone for distinguishing newborns from adults, they uncover complementary, non-redundant biological signatures from the oral cavity that would otherwise be missed by conventional accuracy metrics.

Ruthbah, C. A., Sadi, T. H., Jahan, N. E. S., Adib, A. N. M. T.2026-09-15
💻 bioinformatics

Hierarchical temporal transformer for cancer grade prediction and cross cancer transfer learning from pathology reports

This paper introduces the Hierarchical Temporal Transformer (HTT), a two-level architecture that leverages longitudinal pathology reports and cancer-type embeddings to achieve state-of-the-art cancer grade prediction and demonstrate effective zero-shot transfer learning to unseen cancer types without performance penalties.

Brimo, N., Anand, R., Harb, H., Serdaroglu, D. C.2026-09-15
💻 bioinformatics

POME: Graph-based embeddings for partially observed mixed-type data

The paper introduces POME, a self-supervised graph-based embedding model designed to generate low-dimensional representations for partially observed mixed-type biomedical data, achieving state-of-the-art imputation performance and enabling effective downstream tasks such as patient subgroup discovery, predictive modeling, and therapy recommendation.

Woller, F., Arend, L., Kist, A. M., List, M., Rahimi, F., Sirocchi, C., Blumenthal, D. B.2026-09-15
💻 bioinformatics

Matched full-UDG and non-UDG ancient DNA libraries reveal trade-offs in post-mortem damage correction for imputation and kinship inference

By comparing matched full-UDG and non-UDG ancient DNA libraries from medieval Mongolian individuals, this study demonstrates that while various computational damage correction methods (trimming, rescaling, and masking) differentially balance site retention against error reduction, the optimal choice depends on the specific downstream analysis, with corrected data proving essential for accurate kinship inference in non-UDG libraries.

Ravdandorj, O., Sampildondov, C., Janchiv, K., Gakuhari, T.2026-09-13
💻 bioinformatics

A Framework for Quantifying DNA Methylation Heterogeneity and Detecting Co-methylated loci from Native Nanopore Sequencing

This paper presents a scalable framework integrated into the DMRcaller package that leverages native Oxford Nanopore sequencing to quantify single-molecule DNA methylation heterogeneity and detect co-methylated loci, revealing epigenetic patterns and regulatory interactions obscured by traditional site-level average analyses.

Kim, Y. J., Zabet, N. R.2026-09-13