Bioinformatics sits at the exciting intersection where biology meets data science, using powerful computer tools to decode the vast complexity of living systems. From mapping the human genome to tracking how viruses evolve, this field transforms raw biological information into actionable insights that drive modern medicine and research forward without requiring a supercomputer to understand the basics.

On Gist.Science, we ensure you never miss a breakthrough by processing every new preprint in this category directly from bioRxiv. Our team provides both plain-language explanations and detailed technical summaries for each paper, making cutting-edge discoveries accessible to everyone regardless of their background.

Below are the latest bioinformatics papers added from bioRxiv, ready for you to explore with clarity and depth.

💻 bioinformatics

Scanning transcriptomes for nonlinear, domain-level similarities using hmSEEKR

The paper introduces hmSEEKR, a k-mer-based hidden Markov model that scans transcriptomes to identify non-linear, domain-level sequence similarities in long noncoding RNAs, thereby enabling the discovery of functionally related RNA domains and their associated protein interaction networks without requiring prior knowledge of sequence alignment.

Li, S., Sprague, D. A., Eberhard, Q. E., Boyson, S. P., Laederach, A., Calabrese, J. M.2026-07-08
💻 bioinformatics

AllTheBacteria: a community resource empowers biology and discovers novel peptide antibiotics

The paper introduces AllTheBacteria, a comprehensive community resource that uniformly processes nearly 2.5 million public bacterial genomes to enable large-scale biological research and successfully demonstrates its utility by discovering and validating novel peptide antibiotics through AI-driven mining.

Hunt, M., Torres, M. D. T., Alikhan, N.-F., Anderson, D., Andreani, M. L., Blom, J., Bouras, G., Brinkman, F., Carroll (…)2026-07-07
💻 bioinformatics

PHI: A Galaxy-based workflow for reproducible prophage-host interaction analysis and standardized viral-genomics reporting

The paper introduces PHI, a user-friendly and automated Galaxy-based workflow that streamlines reproducible prophage identification, host interaction prediction, and functional characterization while generating standardized reports to lower barriers for both experts and non-experts in viral genomics research.

Saraiva, J. P., Borim Correa, F., Bernt, M., Ghanem, N., Nieto, E., Brizola Toscan, R., Y. Wick, L., Chatzinotas, A.2026-07-07
💻 bioinformatics

Scop3P in 2026: an expanded proteomics-informed resource contextualizing phosphorylation sites through sequence, structure, mutation, and experimental provenance

The 2026 update to Scop3P presents a major expansion of this proteomics-informed knowledgebase by integrating 152,350 uniformly reprocessed human phosphorylation sites with comprehensive structural, biophysical, evolutionary, and mutational annotations to provide a scalable, provenance-aware resource for interpreting phosphorylation in functional and disease contexts.

Ramasamy, P., Tichshenko, N., Diaz, A., Velghe, K., Massignani, E., Vranken, W. F., Martens, L.2026-07-06
💻 bioinformatics

Benchmarking AlphaFold and related deep learning approaches for modeling antibody and TCR antigen recognition

This study evaluates the performance of AlphaFold2, AlphaFold3, and enhanced sampling protocols in modeling antibody and TCR antigen recognition, demonstrating that while increased sampling and AlphaFold3 generally improve accuracy, success varies significantly across complex types and can be further boosted by pooling complementary models.

Yin, R., Saravanakumar, S., Shi, S. Y., Park, M., Lin, V., Lee, J., Cheung, M., Felbinger, N., Kaufman, S., Eisenberg, M (…)2026-07-06
💻 bioinformatics

Selecting Chromosomes for Polygenic Traits: Algorithms and Complexity

This paper defines and analyzes the NP-complete problem of selecting genomic blocks from multiple source genomes to optimize polygenic traits, proposing a suite of algorithms—including a certified Branch-and-Bound solver, a fast Block-Coordinate-Descent heuristic, and a semidefinite-programming relaxation—that collectively provide optimal or near-optimal solutions with theoretical guarantees and empirical validation on yeast-scale simulations.

Zuk, O.2026-07-05
💻 bioinformatics

Improving Generalizability in Whole-Cell Antibiotic Discovery Through Active Learning

This study demonstrates that a calibrated active learning strategy, optimized through retrospective simulations and validated in a closed-loop *Borrelia burgdorferi* screening campaign, significantly enhances experimental hit rates and enables the training of generalizable machine learning models capable of accurately predicting antibiotic activity for out-of-distribution compounds.

Serrano, L. R., Zhou, A., Wei, Z., Stocks, K.-L. K., Ektefaie, Y., Gwynne, P. J., Chen, E., Krieger, I., Sacchettini, J. (…)2026-07-05