Bioinformatics sits at the exciting intersection where biology meets data science, using powerful computer tools to decode the vast complexity of living systems. From mapping the human genome to tracking how viruses evolve, this field transforms raw biological information into actionable insights that drive modern medicine and research forward without requiring a supercomputer to understand the basics.

On Gist.Science, we ensure you never miss a breakthrough by processing every new preprint in this category directly from bioRxiv. Our team provides both plain-language explanations and detailed technical summaries for each paper, making cutting-edge discoveries accessible to everyone regardless of their background.

Below are the latest bioinformatics papers added from bioRxiv, ready for you to explore with clarity and depth.

💻 bioinformatics

CodonMamba: a foundation model for programmable mRNA coding sequence design

CodonMamba is a state-of-the-art foundation model for mRNA coding sequence design that, through pretraining on large-scale corpora, achieves superior performance in prediction tasks and enables programmable, inference-time steering of codon optimization to meet specific host or application preferences without retraining.

Lang, M., Fang, X., Wang, Z., Chen, M., Cheng, Z., Zhu, X., Tam, K. Y., Zhang, J., Li, X.2026-08-24
💻 bioinformatics

Delta Marches: Generative AI based image synthesis to decode disease-driving morphologic transformations.

Delta-Marches is a generative AI framework that decodes disease mechanisms by simulating idealized morphological transitions between tissue classes to pinpoint subcellular features driving pathophysiological changes, as demonstrated in renal carcinoma grading and colorectal dysplasia.

Nguyen, T. H., Panwar, V., Jarmale, V., Perny, A., Dusek, C., Cai, Q., Kapur, P. H., Danuser, G., Rajaram, S.2026-08-22
💻 bioinformatics

KRAKEN: A provenance-tracked knowledge graph for multiomic and wellness research

KRAKEN is a scalable, provenance-tracked knowledge graph that addresses the underrepresentation of multiomic and wellness data by integrating diverse biomedical sources into a modular, 15-million-node system featuring standardized semantics, built-in analytical tools, and multi-modal interfaces for both human and AI-driven research.

Glen, A. K., Witherington, D., Leslie, T., Baumgartner, A., Fernando, A., Vemuri, B., Nahman, O., Glusman, G., Hood, L. (…)2026-08-22
💻 bioinformatics

Sequence-Derived Representations versus Pfam-Domain Content for Biosynthetic Gene Cluster Retrieval

This study demonstrates that explicit Pfam-domain content outperforms or matches ESM-2 sequence-derived representations for retrieving biosynthetic gene clusters, indicating that sequence embeddings do not currently improve the recovery of alternative biosynthetic pathways beyond established domain-based metrics.

Urokov, R., Khan, A., Eshboyev, F., Asadov, D., Rahman, S., Kushokova, D.2026-08-22
💻 bioinformatics

De novo Design of Macrocyclic Molecular Glues

This paper introduces EvoBind-multimer, a deep learning framework that enables the de novo design of macrocyclic molecular glues from protein sequences alone to induce proximity and drive targeted protein degradation, while also revealing that the functional outcome of these glues can vary context-dependently between different patient-derived models.

Brunner, A., Wierbilowicz, K., Daumiller, D., Bexell, D., Karlsson, K., Sangfelt, O., Bryant, P.2026-08-22
💻 bioinformatics

A novel benchmark dataset for enzyme function prediction reveals the limitations of state-of-the-art models

The paper introduces EnzymARC, a novel benchmark dataset of structurally disrupted enzyme decoys, to demonstrate that current state-of-the-art enzyme function predictors rely heavily on phylogenetic shortcuts and fail to distinguish catalytically incompetent variants from functional enzymes, thereby highlighting the critical need for structure-aware negative examples in model training and evaluation.

Sartori, J., Guimaraes, A. C. R., Machado, L. d. A.2026-08-22
💻 bioinformatics

antigen-prime: Simulating coupled genetic and antigenic evolution of influenza virus

The paper introduces antigen-prime, a forward-time epidemic simulator that links genetic sequences to antigenic phenotypes under host selection to generate ground-truth data for benchmarking influenza variant assignment and growth rate estimation methods, revealing both their accuracy and specific failure modes.

Thornton, Z. T., Tran, T., Figgins, M. D., Huddleston, J., Bedford, T., Matsen, F. A., Haddox, H. K.2026-08-21
💻 bioinformatics

Degree-ranked gene lists omit the cross-module connectors, and a partition-free centrality recovers them

This study reveals that standard degree-based gene prioritization systematically overlooks non-hub connector genes essential for coordinating biological processes, and proposes EDVS, a partition-free, information-theoretic centrality measure that successfully recovers these omitted genes without relying on functional annotations or community detection.

Qun, Z., Huaizheng, Z., Yuxin, Z., Jieying, B., Tan, S.2026-08-21