Bioinformatics sits at the exciting intersection where biology meets data science, using powerful computer tools to decode the vast complexity of living systems. From mapping the human genome to tracking how viruses evolve, this field transforms raw biological information into actionable insights that drive modern medicine and research forward without requiring a supercomputer to understand the basics.

On Gist.Science, we ensure you never miss a breakthrough by processing every new preprint in this category directly from bioRxiv. Our team provides both plain-language explanations and detailed technical summaries for each paper, making cutting-edge discoveries accessible to everyone regardless of their background.

Below are the latest bioinformatics papers added from bioRxiv, ready for you to explore with clarity and depth.

💻 bioinformatics

AnnoAudit: a marker-based protocol for auditing single-cell atlas annotations reveals systematic, state-dependent annotation failure in a widely used traumatic brain injury resource

This paper introduces AnnoAudit, a marker-based protocol that reveals systematic annotation errors in widely used single-cell atlases, such as the CEREBRI traumatic brain injury resource, where the majority of cells labeled as glutamatergic neurons are actually non-neuronal, leading to significant biological artifacts that are corrected upon re-annotation.

Zhang L, Yan Q, Rao H, Li M, Qian X, Zhang Y, Gao R2026-09-08✓ Author reviewed
💻 bioinformatics

End-to-end plaque counting and virus titration from laboratory plate images with deep learning

This paper introduces Titra, an end-to-end deep learning workflow that automates the entire virus titration process—from well detection and plaque segmentation to PFU/mL estimation—demonstrating strong agreement with manual annotations across diverse viral species and plate formats.

Moris, E., Costable, A., Rey, S., Ferreiro, I., Hurtado, J., Villagran, M., Luciano, L. L., Vazquez, A. E., Ramos, J., M (…)2026-09-08
💻 bioinformatics

How do Co-Folding Models Organize Structural Information?

This paper dissects the Boltz-1 co-folding model to reveal that structural information is organized across three distinct streams—single, intra-chain, and inter-chain representations—that follow a Mix-Compress-Refine trajectory, where intra-chain geometry is largely pre-conditioned while inter-chain arrangements are progressively constructed and reconciled by the diffusion module.

Park, M., Kim, S., Moon, S., Kim, H., Jeon, G., Kim, W. Y.2026-09-08
💻 bioinformatics

Keloid transcriptomics reveal heterogeneity in fibroblast subtype enrichment, gene expression, and immune cell responses

This study utilizes bulk RNA-Seq analysis of keloid and matched normal skin tissues to demonstrate that accounting for cell type heterogeneity reveals distinct differences in fibroblast and immune cell enrichment, their specific interactions, and key gene expression signatures underlying keloid disease.

Panzer, J. J., Pan, M., Nair, M., Loveless, I. M., Adrianto, I., Huang, L., Chitale, D., Francescone, R., Vendramini-Cos (…)2026-09-07
💻 bioinformatics

Trustworthy ML/AI for Aging Clocks: Preventing Systematic Prediction Bias in Biological Age Estimation

This paper identifies and addresses the critical issue of systematic prediction bias in machine learning-based aging clocks, which can distort downstream association analyses, by proposing a constrained optimization framework to ensure valid biological age estimation and inference.

Lee, H., Ye, Z., Yang, Y., Pan, Y., Maron, B., Wang, Z., Kochunov, P., Thompson, P., Hong, L. E., MA, T., Chen, C., Chen (…)2026-09-06
💻 bioinformatics

RevPert: predicting candidate drivers of transcriptomic state transitions via gallery-native reverse perturbation

RevPert is a gallery-native reverse perturbation model that prioritizes candidate genetic drivers of transcriptomic state transitions by combining signed Pearson connectivity with a learned residual, demonstrating superior performance in recovering held-out interventions and identifying disease-relevant anchors compared to existing baselines.

Liang, S., Yang, C., Wang, J., Li, y.2026-09-06
💻 bioinformatics

Coarse composition suffices: tabular in-context learning for multi-activity antimicrobial peptide profiling

This paper demonstrates that a simple, sequence-only pipeline combining 330 interpretable descriptors with the TabPFN foundation model outperforms complex, structure-conditioned deep learning approaches in multi-label antimicrobial peptide profiling on the ESCAPE benchmark, achieving state-of-the-art accuracy without the need for gradient-based training or structural data.

Pal, A., Kumar, R., Solanki, D., Pareek, P., Singh, J., Singla, J.2026-09-06