Bioinformatics sits at the exciting intersection where biology meets data science, using powerful computer tools to decode the vast complexity of living systems. From mapping the human genome to tracking how viruses evolve, this field transforms raw biological information into actionable insights that drive modern medicine and research forward without requiring a supercomputer to understand the basics.

On Gist.Science, we ensure you never miss a breakthrough by processing every new preprint in this category directly from bioRxiv. Our team provides both plain-language explanations and detailed technical summaries for each paper, making cutting-edge discoveries accessible to everyone regardless of their background.

Below are the latest bioinformatics papers added from bioRxiv, ready for you to explore with clarity and depth.

💻 bioinformatics

Predicting optimal growth temperatures of bacteria using learned structural information from a single protein

The paper introduces ROSEATE, a novel framework that accurately predicts bacterial optimal growth temperatures by leveraging MSA Transformer-derived structural signatures from a single ubiquitous protein, adenylate kinase, enabling robust, phylogenetically generalizable, and community-level thermal inference across diverse environments.

Hoffert, M., Myerscough, D., Dragone, N. B., Gebert, M. J., Silberg, J. J., Fierer, N.2026-06-18
💻 bioinformatics

MetaHarmonizer: robust biomedical metadata harmonization and a contamination control for inflated LLM performance on public benchmarks

MetaHarmonizer is a robust, fully local, and deterministic automated system for biomedical metadata harmonization that combines a multi-stage cascade with controlled vocabularies to prevent hallucinations and inflated benchmark performance, achieving state-of-the-art accuracy in schema and ontology mapping while enabling principled human-in-the-loop triage.

Li, C., Dahl, A., Gravel-Pucillo, K. D., Long, K., Waters, M., de Bruijin, I., Davis, S., Oh, S.2026-06-17
💻 bioinformatics

VLab4Mic: prediction of structural resolvability in super-resolution microscopy

VLab4Mic is a simulation platform that predicts the structural resolvability of protein assemblies across various super-resolution microscopy modalities by modeling probe placement and steric constraints, thereby enabling researchers to assess experimental feasibility before conducting physical experiments.

Martinez, D., Saraiva, B. M., Shakespeare, T., Bates, M., Owen, D. M., Leterrier, C., Del Rosario, M., Henriques, R.2026-06-16
💻 bioinformatics

THEOBROMA: an aggregated open database of 1.13 million natural products with per-compound license auditing, three-tier classification, and stereochemistry-aware deduplication

THEOBROMA is an open database aggregating over 1.13 million natural products from 29 sources that distinguishes itself through per-compound license auditing, a three-tier classification system, and stereochemistry-aware deduplication to enable license-compliant virtual screening and isomer-specific bioactivity analysis.

Klamt, T., Jaczkowski, A., Franke, J., Nejdl, W.2026-06-16
💻 bioinformatics

Rapid and consistent clustering of millions of genomes highlights the diversity of prokaryotic life

The authors present gemsparcl, a highly scalable and efficient tool that clusters over 5.6 million bacterial genomes into 92,954 species-level genomic cohesive units in approximately 14 hours, thereby overcoming computational bottlenecks to enable comprehensive, reference-free analysis of prokaryotic diversity and taxonomy.

von Wachsmann, J. H., Lorenz, L. J., Gurbich, T. A., Russell, M. J., Rodriguez Bouza, V., Horsfield, S. T., Lees, J. A. (…)2026-06-15
💻 bioinformatics

WitChi: Efficient Detection and Pruning of Compositional Bias in Phylogenomic Alignments Using Empirical Chi-Squared Testing

WitChi is a computationally efficient tool that uses empirical chi-squared testing to detect and iteratively prune compositionally biased sites from large-scale phylogenomic alignments, thereby restoring accurate phylogenetic topologies without the high computational cost of complex composition-aware models.

Koestlbacher, S., Panagiotou, K., Tamarit, D., Ettema, T.2026-06-13
💻 bioinformatics

Testing the reliability of AI-generated protein structures

This study demonstrates that AlphaFold2 and ColabFold exhibit a very low false positive rate when predicting structures for non-protein sequences, while serendipitously revealing that some high-scoring predictions in noncoding human genomic regions correspond to previously unannotated pseudogenes, suggesting potential errors in existing gene annotations and validating the utility of high-scoring structural predictions for further investigation.

Xu, A., Salzberg, S.2026-06-13
💻 bioinformatics

ADMETron: An AI-driven SaaS platform for comprehensive ADMET prediction and compound prioritisation

ADMETron is an AI-driven SaaS platform that integrates RNN-derived molecular embeddings with gradient boosting machines to predict 34 ADMET endpoints and features an interactive radar graph visualization tool, enabling comprehensive, high-throughput compound prioritization and data-driven decision-making in drug discovery.

Nair, D. N., Yadav, R. S., Jondhale, P. M., Didhate, S., Gunjal, G., Ranjit, A., Patil, P., Dawande, A., Shisode, A., Bh (…)2026-06-13