Bioinformatics sits at the exciting intersection where biology meets data science, using powerful computer tools to decode the vast complexity of living systems. From mapping the human genome to tracking how viruses evolve, this field transforms raw biological information into actionable insights that drive modern medicine and research forward without requiring a supercomputer to understand the basics.

On Gist.Science, we ensure you never miss a breakthrough by processing every new preprint in this category directly from bioRxiv. Our team provides both plain-language explanations and detailed technical summaries for each paper, making cutting-edge discoveries accessible to everyone regardless of their background.

Below are the latest bioinformatics papers added from bioRxiv, ready for you to explore with clarity and depth.

💻 bioinformatics

Comprehensive characterization of V(D)J recombination from long-read transcriptomic data with VDJcraft

The authors present VDJcraft, a novel pipeline specifically designed for accurate V(D)J recombination analysis using long-read transcriptomic data, which outperforms existing methods in gene detection and recombination accuracy while enabling the discovery of novel gene subclasses and disease-associated immune signatures.

Hu, K., Rosenberg, A. F., Song, Y., Fan, C.-H., Peng, Z., Gao, M., Chong, Z.2026-04-05
💻 bioinformatics

Interpretable Deep Learning-Based Multi-Omics Integrationfor Prognosis in Hepatocellular Carcinoma

This study presents an interpretable, attention-based deep learning framework that integrates mRNA, miRNA, and DNA methylation data to significantly improve prognostic accuracy for hepatocellular carcinoma patients compared to existing models, while identifying biologically relevant biomarkers and demonstrating robust performance on external validation cohorts.

Znabu, B. F., Atif, Z.2026-04-05
💻 bioinformatics

PanTEon: a cross-kingdom framework to guide the design of transposable element classifiers

The paper introduces PanTEon, a cross-kingdom deep learning framework comprising a harmonized database and a modular benchmarking platform that enables reproducible, standardized training and evaluation of transposable element classifiers across diverse eukaryotic lineages.

Orozco-Arias, S., Ferrer-Pomer, I., Rodrigues de Goes, F., Gaviria-Orrego, S., Gomiz-Fernandez, J., Llatser-Torres, J. (…)2026-04-04
💻 bioinformatics

Improved quantitation in data-independent acquisition proteomics via retention time boundary imputation

This paper introduces Nettle, an open-source tool that improves quantitation in data-independent acquisition proteomics by imputing peptide retention time boundaries to integrate chromatographic signals, thereby reducing missing data limitations and enhancing accuracy compared to traditional methods.

Harris, L. J., Riffle, M., Shulman, N., Fondrie, W. E., Wu, C. C., Johnson Erickson, D. P., Morimoto, A., Shaver, B., St (…)2026-04-03
💻 bioinformatics

CellWHISPER disentangles direct cell-cell communication from structural proximity

CellWHISPER is a statistically robust and computationally scalable framework that accurately infers direct, contact-mediated cell-cell communication from spatial transcriptomics data by disentangling true signaling interactions from structural proximity, enabling the discovery of tissue- and disease-specific signaling programs such as gap-junction coupling in the brain.

Kumar, A., Moctezuma, F. R., Aggarwal, B., Zhang, N., Coskun, A. F., Sinha, S.2026-04-03
💻 bioinformatics

GATSBI: Improving context-aware protein embeddingsthrough biologically motivated data splits

The paper introduces GATSBI, a graph attention-based framework that generates context-aware protein embeddings by integrating diverse biological data and employing task-aligned evaluation protocols, demonstrating superior generalization—particularly for understudied proteins—compared to existing methods that rely on biologically inappropriate data splits.

Nayar, G., Altman, R. B.2026-04-03