Bioinformatics sits at the exciting intersection where biology meets data science, using powerful computer tools to decode the vast complexity of living systems. From mapping the human genome to tracking how viruses evolve, this field transforms raw biological information into actionable insights that drive modern medicine and research forward without requiring a supercomputer to understand the basics.

On Gist.Science, we ensure you never miss a breakthrough by processing every new preprint in this category directly from bioRxiv. Our team provides both plain-language explanations and detailed technical summaries for each paper, making cutting-edge discoveries accessible to everyone regardless of their background.

Below are the latest bioinformatics papers added from bioRxiv, ready for you to explore with clarity and depth.

💻 bioinformatics

Tox21mer, A transformer foundation model for Tox21 high-throughput concentration-response curves data

The paper introduces Tox21mer, a 43.5-million-parameter transformer foundation model pretrained on 2.5 million Tox21 concentration-response curves via masked-response reconstruction, which generates high-quality 768-dimensional embeddings that achieve state-of-the-art performance in predicting assay outcomes and AC50 values while enabling extrapolation to untested compounds.

Li, L., Hwang, J., Shockley, K., Li, Y., Motsinger-Reif, A., Hsieh, J.-H., Auerbach, S. S., Reif, D.2026-06-19
💻 bioinformatics

Children's DNA Methylation and Family Dynamics in a Congo Basin Subsistence Community: Links with Parental Conflict and Fathers' Caregiving

This study demonstrates that in a Congo Basin subsistence community, children's DNA methylation patterns are significantly associated with parental conflict and father caregiving, linking these family dynamics to genes involved in stress, immunity, and development, thereby suggesting that the biological embedding of family environments is a universal phenomenon across diverse socio-ecological contexts.

Chan, M. H.-M., Merrill, S. S., Zhuang, B. C., Lin, D. T. S., Macisaac, J. L., Miegakanda, V., Lew-Levy, S., Boyette, A. (…)2026-06-19
💻 bioinformatics

Predicting optimal growth temperatures of bacteria using learned structural information from a single protein

The paper introduces ROSEATE, a novel framework that accurately predicts bacterial optimal growth temperatures by leveraging MSA Transformer-derived structural signatures from a single ubiquitous protein, adenylate kinase, enabling robust, phylogenetically generalizable, and community-level thermal inference across diverse environments.

Hoffert, M., Myerscough, D., Dragone, N. B., Gebert, M. J., Silberg, J. J., Fierer, N.2026-06-18
💻 bioinformatics

MetaHarmonizer: robust biomedical metadata harmonization and a contamination control for inflated LLM performance on public benchmarks

MetaHarmonizer is a robust, fully local, and deterministic automated system for biomedical metadata harmonization that combines a multi-stage cascade with controlled vocabularies to prevent hallucinations and inflated benchmark performance, achieving state-of-the-art accuracy in schema and ontology mapping while enabling principled human-in-the-loop triage.

Li, C., Dahl, A., Gravel-Pucillo, K. D., Long, K., Waters, M., de Bruijin, I., Davis, S., Oh, S.2026-06-17
💻 bioinformatics

VLab4Mic: prediction of structural resolvability in super-resolution microscopy

VLab4Mic is a simulation platform that predicts the structural resolvability of protein assemblies across various super-resolution microscopy modalities by modeling probe placement and steric constraints, thereby enabling researchers to assess experimental feasibility before conducting physical experiments.

Martinez, D., Saraiva, B. M., Shakespeare, T., Bates, M., Owen, D. M., Leterrier, C., Del Rosario, M., Henriques, R.2026-06-16
💻 bioinformatics

THEOBROMA: an aggregated open database of 1.13 million natural products with per-compound license auditing, three-tier classification, and stereochemistry-aware deduplication

THEOBROMA is an open database aggregating over 1.13 million natural products from 29 sources that distinguishes itself through per-compound license auditing, a three-tier classification system, and stereochemistry-aware deduplication to enable license-compliant virtual screening and isomer-specific bioactivity analysis.

Klamt, T., Jaczkowski, A., Franke, J., Nejdl, W.2026-06-16
💻 bioinformatics

Rapid and consistent clustering of millions of genomes highlights the diversity of prokaryotic life

The authors present gemsparcl, a highly scalable and efficient tool that clusters over 5.6 million bacterial genomes into 92,954 species-level genomic cohesive units in approximately 14 hours, thereby overcoming computational bottlenecks to enable comprehensive, reference-free analysis of prokaryotic diversity and taxonomy.

von Wachsmann, J. H., Lorenz, L. J., Gurbich, T. A., Russell, M. J., Rodriguez Bouza, V., Horsfield, S. T., Lees, J. A. (…)2026-06-15
💻 bioinformatics

WitChi: Efficient Detection and Pruning of Compositional Bias in Phylogenomic Alignments Using Empirical Chi-Squared Testing

WitChi is a computationally efficient tool that uses empirical chi-squared testing to detect and iteratively prune compositionally biased sites from large-scale phylogenomic alignments, thereby restoring accurate phylogenetic topologies without the high computational cost of complex composition-aware models.

Koestlbacher, S., Panagiotou, K., Tamarit, D., Ettema, T.2026-06-13