Bioinformatics sits at the exciting intersection where biology meets data science, using powerful computer tools to decode the vast complexity of living systems. From mapping the human genome to tracking how viruses evolve, this field transforms raw biological information into actionable insights that drive modern medicine and research forward without requiring a supercomputer to understand the basics.

On Gist.Science, we ensure you never miss a breakthrough by processing every new preprint in this category directly from bioRxiv. Our team provides both plain-language explanations and detailed technical summaries for each paper, making cutting-edge discoveries accessible to everyone regardless of their background.

Below are the latest bioinformatics papers added from bioRxiv, ready for you to explore with clarity and depth.

💻 bioinformatics

HoloCell: A Generative Foundation Model for Holistic Cellular Modeling

HoloCell is a 860-million-parameter generative foundation model pretrained on a massive multi-omics corpus that unifies epigenomic, transcriptomic, and proteomic data through hierarchical tokenization and iterative diffusion to enable holistic cellular representation learning and flexible cross-modal generation.

Jiang, Q., Li, Z., Hu, B., Bie, Y., Li, K., Li, Q., Jin, P., He, Y., Deng, P., Wang, Z., Chen, X., Qin, T., Liu, H., Jia (…)2026-06-23
💻 bioinformatics

Comorbidity structure as an inductive bias: Comparing output-head designs for multi-label prediction of diabetes and myocardial infarction complications

This paper demonstrates that output-head designs for multi-label prediction should be explicitly chosen to reflect the underlying biological structure of comorbidities, as evidenced by a symmetric conditional random field outperforming complex alternatives in the microvascularly linked complications of Type 2 diabetes, while no single architecture proved stable for the heterogeneous electrophysiological complications of myocardial infarction.

Asumboya, W. A., Agbenorhevi, P. K., Adams, C. F., Ayariga, D. A., Adjadeh, T., Adams Ziblim, S., Kwofie, S. K.2026-06-23
💻 bioinformatics

Drug-Prot: A query system for statistical inference of drug effects and interactions in dynamic proteomic networks

Drug-Prot is a publicly available computational framework and web application that leverages large-scale perturbation proteomics data from breast cancer cell lines to statistically infer causal drug effects, drug-drug interactions, and dynamic protein dependency networks, thereby enabling targeted analysis of protein-level responses to single and combination therapies.

Ulmer, M., Sun, R., Qian, L., Aebersold, R., Guo, T., Buehlmann, P.2026-06-22
💻 bioinformatics

Hierarchical classification of immune cell transcriptomes at population-scale

This paper introduces Suco, a resource of independent expert annotations, and Compocyte, a hierarchical classifier, to establish a robust framework that successfully classified 15.6 million immune cells across nearly 4,000 patients, revealing novel immune phenotypes and advancing population-scale immunology research.

Beltz, C., Qiu, Z., Sadowski, L., Kraske, J. A., Aggarwal, A., Quintanal-Villalonga, A., Manoj, P., Littbarski, A., Baja (…)2026-06-21
💻 bioinformatics

Antibody-Antigen Affinity Prediction with Chain-Aware Protein Language Modeling

The paper introduces AbAffinity, a lightweight, sequence-only deep learning model that utilizes a chain-aware three-stream architecture to accurately predict antibody-antigen affinity by preserving distinct heavy chain, light chain, and antigen representations, thereby outperforming existing methods in scenarios where structural data is unavailable.

Singh, H., Malhotra, A., Srivastava, S. P., SINGH, R. K., Gorantla, R.2026-06-21
💻 bioinformatics

The recount3 Python package for programmatic access to uniformly processed RNA-seq data

The recount3 Python package provides a robust API and CLI to enable efficient programmatic access, caching, and analysis-ready data formatting for tens of thousands of uniformly processed human and mouse RNA-seq samples, thereby bridging the gap between large-scale public transcriptomic data and modern Python-based machine learning ecosystems.

Alsalihi, A., Flight, R. M., Moseley, H. N. B.2026-06-20
💻 bioinformatics

Tox21mer, A transformer foundation model for Tox21 high-throughput concentration-response curves data

The paper introduces Tox21mer, a 43.5-million-parameter transformer foundation model pretrained on 2.5 million Tox21 concentration-response curves via masked-response reconstruction, which generates high-quality 768-dimensional embeddings that achieve state-of-the-art performance in predicting assay outcomes and AC50 values while enabling extrapolation to untested compounds.

Li, L., Hwang, J., Shockley, K., Li, Y., Motsinger-Reif, A., Hsieh, J.-H., Auerbach, S. S., Reif, D.2026-06-19
💻 bioinformatics

Children's DNA Methylation and Family Dynamics in a Congo Basin Subsistence Community: Links with Parental Conflict and Fathers' Caregiving

This study demonstrates that in a Congo Basin subsistence community, children's DNA methylation patterns are significantly associated with parental conflict and father caregiving, linking these family dynamics to genes involved in stress, immunity, and development, thereby suggesting that the biological embedding of family environments is a universal phenomenon across diverse socio-ecological contexts.

Chan, M. H.-M., Merrill, S. S., Zhuang, B. C., Lin, D. T. S., Macisaac, J. L., Miegakanda, V., Lew-Levy, S., Boyette, A. (…)2026-06-19