Bioinformatics sits at the exciting intersection where biology meets data science, using powerful computer tools to decode the vast complexity of living systems. From mapping the human genome to tracking how viruses evolve, this field transforms raw biological information into actionable insights that drive modern medicine and research forward without requiring a supercomputer to understand the basics.

On Gist.Science, we ensure you never miss a breakthrough by processing every new preprint in this category directly from bioRxiv. Our team provides both plain-language explanations and detailed technical summaries for each paper, making cutting-edge discoveries accessible to everyone regardless of their background.

Below are the latest bioinformatics papers added from bioRxiv, ready for you to explore with clarity and depth.

💻 bioinformatics

Development of Deep-Learning Models that Predict Quantitative Protein-Ligand Interac-tions in Glycobiology as a part of a Capstone Course

As part of a University of Alberta capstone course, this paper introduces three deep-learning models (ProMax, APEX, and UltraMax) trained on a hybrid dataset of approximately one million protein-ligand pairs to predict quantitative glycan-protein binding strengths, while highlighting the challenges posed by long-tail data distributions and insufficient chiral feature utilization.

Yin, H., Liu, W., Zhou, W., Chang, Z., Carpenter, E. J., Satyajith, A., Haregu, S., Greiner, R., Derda, R.2026-06-24
💻 bioinformatics

ComplexDesign: sequence-hallucination design of protein binders bridging multiple proteins

ComplexDesign is a hallucination-based approach that utilizes structure-prediction-guided sequence optimization and a specialized masking mechanism to successfully design multichain protein complexes and flexible binders bridging multiple targets, outperforming existing methods in both unconditional multimer design and ternary complex generation.

Xu, J., Ren, M., Qi, N., Zhang, X., He, Z., Yu, C., Bu, D.2026-06-24
💻 bioinformatics

HoloCell: A Generative Foundation Model for Holistic Cellular Modeling

HoloCell is a 860-million-parameter generative foundation model pretrained on a massive multi-omics corpus that unifies epigenomic, transcriptomic, and proteomic data through hierarchical tokenization and iterative diffusion to enable holistic cellular representation learning and flexible cross-modal generation.

Jiang, Q., Li, Z., Hu, B., Bie, Y., Li, K., Li, Q., Jin, P., He, Y., Deng, P., Wang, Z., Chen, X., Qin, T., Liu, H., Jia (…)2026-06-23
💻 bioinformatics

Comorbidity structure as an inductive bias: Comparing output-head designs for multi-label prediction of diabetes and myocardial infarction complications

This paper demonstrates that output-head designs for multi-label prediction should be explicitly chosen to reflect the underlying biological structure of comorbidities, as evidenced by a symmetric conditional random field outperforming complex alternatives in the microvascularly linked complications of Type 2 diabetes, while no single architecture proved stable for the heterogeneous electrophysiological complications of myocardial infarction.

Asumboya, W. A., Agbenorhevi, P. K., Adams, C. F., Ayariga, D. A., Adjadeh, T., Adams Ziblim, S., Kwofie, S. K.2026-06-23
💻 bioinformatics

Drug-Prot: A query system for statistical inference of drug effects and interactions in dynamic proteomic networks

Drug-Prot is a publicly available computational framework and web application that leverages large-scale perturbation proteomics data from breast cancer cell lines to statistically infer causal drug effects, drug-drug interactions, and dynamic protein dependency networks, thereby enabling targeted analysis of protein-level responses to single and combination therapies.

Ulmer, M., Sun, R., Qian, L., Aebersold, R., Guo, T., Buehlmann, P.2026-06-22
💻 bioinformatics

Hierarchical classification of immune cell transcriptomes at population-scale

This paper introduces Suco, a resource of independent expert annotations, and Compocyte, a hierarchical classifier, to establish a robust framework that successfully classified 15.6 million immune cells across nearly 4,000 patients, revealing novel immune phenotypes and advancing population-scale immunology research.

Beltz, C., Qiu, Z., Sadowski, L., Kraske, J. A., Aggarwal, A., Quintanal-Villalonga, A., Manoj, P., Littbarski, A., Baja (…)2026-06-21
💻 bioinformatics

Antibody-Antigen Affinity Prediction with Chain-Aware Protein Language Modeling

The paper introduces AbAffinity, a lightweight, sequence-only deep learning model that utilizes a chain-aware three-stream architecture to accurately predict antibody-antigen affinity by preserving distinct heavy chain, light chain, and antigen representations, thereby outperforming existing methods in scenarios where structural data is unavailable.

Singh, H., Malhotra, A., Srivastava, S. P., SINGH, R. K., Gorantla, R.2026-06-21
💻 bioinformatics

The recount3 Python package for programmatic access to uniformly processed RNA-seq data

The recount3 Python package provides a robust API and CLI to enable efficient programmatic access, caching, and analysis-ready data formatting for tens of thousands of uniformly processed human and mouse RNA-seq samples, thereby bridging the gap between large-scale public transcriptomic data and modern Python-based machine learning ecosystems.

Alsalihi, A., Flight, R. M., Moseley, H. N. B.2026-06-20