Bioinformatics sits at the exciting intersection where biology meets data science, using powerful computer tools to decode the vast complexity of living systems. From mapping the human genome to tracking how viruses evolve, this field transforms raw biological information into actionable insights that drive modern medicine and research forward without requiring a supercomputer to understand the basics.

On Gist.Science, we ensure you never miss a breakthrough by processing every new preprint in this category directly from bioRxiv. Our team provides both plain-language explanations and detailed technical summaries for each paper, making cutting-edge discoveries accessible to everyone regardless of their background.

Below are the latest bioinformatics papers added from bioRxiv, ready for you to explore with clarity and depth.

💻 bioinformatics

PKProbDesign: RNA inverse folding including pseudoknots by optimizing thermodynamic folding probability

PKProbDesign is a novel sampling-based framework that addresses the challenge of RNA inverse folding for pseudoknotted structures by directly optimizing thermodynamic folding probabilities through scaffold decomposition and conditional ensemble evaluation, outperforming existing methods like DesiRNA and MODENA on benchmark targets.

Otagaki, T., Iwakiri, J., terai, g., Asai, K., Sato, K.2026-07-11
💻 bioinformatics

Spatial mitochondrial lineage tracing uncovers a premetastatic niche and microenvironment programmed fate switching in osteosarcoma

By integrating single-cell and spatial transcriptomics with mitochondrial variant-based lineage tracing, this study reveals that osteosarcoma progression involves a stepwise transcriptional cascade from COL3A1 progenitors to metastatic THY1 cells, a fate switch strictly licensed by spatial co-localization with dysfunctional PLVAP endothelia to form a PTN/NOTCH-enriched premetastatic niche, thereby establishing a "time-space-lineage" framework where microenvironmental signals drive irreversible clonal commitment.

Xue, Y., Su, Z., Su, J., Chu, A. S., Wan, L., Chen, M., Cheung, J. P. Y., Cheung, K. S. C., Ho, J. W. K.2026-07-11
💻 bioinformatics

ProtAug: An Empirical Investigation of pLM-Guided Data Augmentation for Protein Sequence Prediction Tasks

This paper introduces ProtAug, a framework for pLM-guided protein sequence data augmentation that systematically demonstrates how user-controlled variation levels can consistently improve downstream prediction performance across diverse tasks, revealing that while preserving biological plausibility is beneficial at low-to-moderate variation, high-variation augmentation can also drive gains through regularization effects.

Chen, Z., Wang, R., Luo, Q.2026-07-11
💻 bioinformatics

MKMC enables reference-free transcriptomic analysis using k-mer representations

MKMC is a scalable, reference-free toolkit that leverages k-mer statistics to enable robust RNA-seq analysis across model and non-model organisms, successfully detecting biological signals and isoform-specific events that often elude traditional alignment-based methods.

Mboning, L., Dlugosz, M., Kokot, M., Chen, J., Costa, E. K., Wu, M.-R., Wang, S., Bouchard, L.-S., Deorowicz, S., Pelleg (…)2026-07-10