Bioinformatics sits at the exciting intersection where biology meets data science, using powerful computer tools to decode the vast complexity of living systems. From mapping the human genome to tracking how viruses evolve, this field transforms raw biological information into actionable insights that drive modern medicine and research forward without requiring a supercomputer to understand the basics.

On Gist.Science, we ensure you never miss a breakthrough by processing every new preprint in this category directly from bioRxiv. Our team provides both plain-language explanations and detailed technical summaries for each paper, making cutting-edge discoveries accessible to everyone regardless of their background.

Below are the latest bioinformatics papers added from bioRxiv, ready for you to explore with clarity and depth.

💻 bioinformatics

MAAMOUL: Metabolic network-based discovery of microbiome-metabolome shifts in disease

The paper introduces MAAMOUL, a knowledge-based computational framework that integrates metagenomic and metabolomic data via a global metabolic network to identify disease-specific microbial metabolic modules, successfully uncovering coherent functional shifts in inflammatory bowel disease and irritable bowel syndrome that were missed by conventional analytical methods.

Muller, E., Baum, S., Borenstein, E.2026-03-30
💻 bioinformatics

DualLoc: Full-parameter fine-tuning of cascaded dual transformers for protein subcellular localization prediction

DualLoc is a novel multi-label predictor that employs full-parameter fine-tuning of a cascaded dual-transformer architecture to achieve state-of-the-art accuracy in predicting protein subcellular localization across ten compartments, effectively capturing complex multi-compartment localization patterns and biologically relevant organelle couplings.

Chen, Y. G., Chung, W.-Y., Chang, K. Y.2026-03-30
💻 bioinformatics

Clinical evidence yield as a framework for evaluating computational predictors and multiplexed assays of variant effect

This paper introduces mean evidence strength (MES), a novel metric based on ACMG/AMP guidelines and Bayesian calibration, to evaluate computational predictors and multiplexed assays of variant effect by quantifying their clinical evidence yield rather than relying solely on traditional discrimination metrics like AUROC.

Shang, Y., Badonyi, M., Marsh, J. A.2026-03-30
💻 bioinformatics

ALPINE: A Scalable Pipeline for Comprehensive Classification of Gene-Editing Outcomes from Long-Read Amplicon Sequencing

ALPINE is a scalable, reproducible Python-based pipeline that leverages long-read amplicon sequencing to comprehensively classify and quantify diverse gene-editing outcomes, including complex viral vector integrations and structural variants, addressing key limitations of existing short-read tools for therapeutic development.

Chen, Y., Gao, X.-H., Vichas, A., Wang, J., Golhar, R., Neuhaus, I.2026-03-30
💻 bioinformatics

Deciphering sepsis molecular subtypes using large-scale data to identify subtype-specific drug repurposing

By constructing a transcriptomic atlas of 3,713 samples to identify four distinct molecular subtypes of sepsis, this study elucidates subtype-specific pathophysiological mechanisms and proposes targeted drug repurposing strategies, such as corticosteroids for C1 and methylene blue for the high-mortality C4 subtype, to advance precision medicine and explain the failure of previous broad-spectrum clinical trials.

Smith, L. A., Augustin, B., Jacob, V., Black, L. P., Bertrand, A., Hopson, C., Cagmat, E., Datta, S., Reddy, S., Guirgis (…)2026-03-30
💻 bioinformatics

Panmap: Scalable phylogeny-guided alignment, genotyping, and placement on pangenomes

Panmap is a scalable tool that utilizes a phylogenetically compressed k-mer index to efficiently align, genotype, and place sequencing reads onto mutation-annotated pangenomes containing millions of genomes, significantly reducing computational time and storage requirements compared to existing methods.

Kramer, A. M., Zhang, A., Ayala, N., de Sanctis, B., Karim, L. M., Hinrichs, A. S., Walia, S., Turakhia, Y., Corbett-Det (…)2026-03-30
💻 bioinformatics

DeepBranchAI: A Novel Cascade Workflow Enabling Accessible 3D Branching Network Segmentation

DeepBranchAI introduces a novel cascade workflow that overcomes the annotation bottleneck in 3D branching network segmentation by iteratively refining sparse labels through a positive feedback loop of random forests and expert input, ultimately enabling the training of a robust, topology-preserving 3D nnU-Net model that achieves high accuracy across diverse biological and medical datasets while significantly reducing manual annotation time.

Maltsev, A. V., Hartnell, L., Ferrucci, L.2026-03-29
💻 bioinformatics

A run-length-compressed skiplist data structure for dynamic GBWTs supports time and space efficient pangenome operations over syncmers

This paper introduces a dynamic, run-length-compressed skiplist data structure for the graph Burrows-Wheeler transform (GBWT) that enables time and space-efficient pangenome operations on syncmer graphs, successfully building a 5.8 GB lossless representation of 92 human genomes in under an hour and supporting rapid sequence matching.

Durbin, R.2026-03-29
💻 bioinformatics

Open-source, Hardware-Independent GPU Acceleration for Scalable Nanopore Basecalling with Slorado and Openfish

This paper introduces Openfish, an open-source GPU-accelerated decoding library, and Slorado, a fully open-source basecalling framework, to provide a hardware-independent, scalable, and accurate alternative to Oxford Nanopore Technologies' proprietary Dorado software, thereby eliminating hardware lock-in and enhancing accessibility for nanopore sequencing.

Wong, B., Singh, G., Javaid, H., Denolf, K., Liyanage, K., Samarakoon, H., Deveson, I. W., Gamaarachchi, H.2026-03-28