🧬 biology

Geometric Representations of Knowledge Inside Biological Large Language Models: an Empirical Analysis

This empirical study reveals that while biological large language models (SCFMs) do encode a real, low-dimensional, and partially linearized geometric structure for biological knowledge, this structure is modest, often nonlinearly entangled, and largely overlaps with what classical expression-based methods like PCA can already recover, falling short of the crisp, regular geometry observed in natural language models.

Olivia Denvis2026-07-22
🧬 biology

An Empirical Comparison of Virtual Cell Models: Perturbation Prediction, Representation, and the Baseline Gap

This paper presents a unified benchmark of eleven virtual cell models, revealing that while deep learning approaches significantly outperform baselines in representation tasks and combinatorial perturbation prediction, they generally fail to surpass simple linear models in predicting unseen single-gene perturbations, suggesting the field has achieved a "virtual microscope" for cell state analysis rather than a true "virtual simulator" for causal perturbation dynamics.

Olivia Denvis2026-07-22
🧬 biology

Findings from Sparse Autoencoders for DNA Sequence Models: Motif Detectors, Reading-Frame Features, and the Scarcity of Regulatory Logic

This study demonstrates that while sparse autoencoders effectively extract monosemantic, biologically interpretable features like sequence motifs and reading frames from DNA foundation models, they reveal a significant scarcity of features encoding complex regulatory logic, suggesting current models capture a genomic dictionary rather than a grammar engine.

Olivia Denvis2026-07-22
🧬 biology

Information Entropy of Biological Data: An Empirical Analysis

This paper presents a unified empirical analysis across six biological data modalities demonstrating that finite-sample bias severely distorts information-theoretic estimates, and argues that reporting entropy and mutual information without specifying the estimator and sample size is indefensible given that bias-corrected methods reveal distinct biological signals and correct systematic underestimation of entropy and overestimation of information content.

Alexander Memming2026-07-22
🧬 biology

Pre-Transformer Models for Longevity Science and Deep Ageing Clocks: An Empirical Analysis

This empirical study demonstrates that for biological-age prediction across diverse ageing cohorts, pre-transformer models (such as penalized regression and gradient boosting) generally outperform or match attention-based transformers in accuracy, data efficiency, cross-cohort transfer, and mortality risk stratification, establishing them as the rational default for longevity science while positioning transformers as a specialized tool only viable with exceptionally large datasets.

John Feng2026-07-22
🧬 biology

Deep Learning for Longevity Biomarker Development: An Empirical Analysis of Model Capacity versus Prediction Target for Blood-Based Aging Clocks

This empirical study demonstrates that for blood-based aging clocks, the choice of prediction target (mortality-derived phenotypic age versus chronological age) overwhelmingly determines biomarker validity and mortality association, rendering increases in deep learning model capacity largely ineffective for improving these critical outcomes.

John Feng2026-07-22
🧬 biology

EMT activates ER-to-Golgi trafficking through upregulation of REEP2 to promote lung cancer progression

This study identifies REEP2 as a critical mediator of EMT-driven lung cancer progression, demonstrating that the ZEB1/miR-183/193a axis upregulates REEP2 to enhance ER-to-Golgi trafficking and the secretion of pro-tumorigenic factors, thereby promoting tumor metastasis and immunosuppression.

Guan-Yu Xiao, Kevin Fulp, Oluwafunminiyi Obaleye, Shike Wang, Xin Liu, Jiang Yu, Jonathan M. Kurie, Jun Xu, William Russ (…)2026-07-22