📄 other

The Ground-Truth Problem in Regulatory Genomics: An Empirical Analysis of Reference-Dependent Method Rankings

This study demonstrates that the ranking of gene regulatory network inference methods is highly sensitive to the choice of reference network, revealing that over half of method comparisons can reverse depending on the biological definition and coverage of the ground truth, thereby arguing that reference selection should be treated as a critical experimental variable rather than a fixed background assumption.

Liu Chen2026-07-27
📄 other

How Much Does a Benchmark Weight Move a Leaderboard? An Empirical Analysis of Single-Cell Integration Rankings

This empirical analysis of the scIB single-cell integration benchmark reveals that while the published ranking is locally stable, the winning method and overall ordering are sensitive to defensible changes in weight priorities and task composition, suggesting that benchmark reports should include sensitivity curves and uncertainty estimates rather than relying solely on a single rank table.

Liu Chen2026-07-27
📄 other

Reading Mechanism off Biological Foundation Models: An Empirical Analysis of Genomic Neighborhood Structure in scGPT Gene Embeddings

This empirical study demonstrates that the scGPT biological foundation model exhibits a small but reproducible association between its static gene embedding geometry and linear genomic neighborhood, revealing that genes located close together on the same chromosome are more similar in the model's representation than distant genes, even though this spatial information was not explicitly provided during training.

Liu Chen2026-07-27
📄 other

The effect of dietary fiber based on fermentability and viscosity on the gut microbiome composition in chronic kidney disease: a systematic review of experimental and clinical trials

This systematic review of 22 experimental and clinical trials concludes that isolated dietary fiber interventions yield variable and largely inconsistent effects on gut microbiota composition in chronic kidney disease patients, with no clear pattern emerging based solely on fiber fermentability or viscosity.

Seyedeh Nooshan Mirmohammadali, Claudia Carrillo, Jason B. Reed, Brandon M. Kistler, Hannah E. Wilson, Bruce Hamaker, Sh (…)2026-07-27
📄 other

Regional sequence and embedding divergence among five grass OsNAC25-like proteins

This study analyzes five grass OsNAC25-like proteins to demonstrate that while sequence identity and structure-prediction confidence consistently show higher conservation in the N-terminal DNA-binding domain compared to the C-terminal region, protein-language-model embeddings are significantly influenced by specific candidate identity, highlighting the finger-millet sequence as a priority for further phylogenetic investigation.

Neil Sumanth2026-07-27
📄 other

Natural Language Processing Psychometrics

This paper introduces "NLP Psychometrics," a framework that leverages Large Language Model personas and interpretable AI techniques to link psychological prediction scores to specific linguistic features, demonstrating that while synthetic data can effectively expose biases and predict mental health outcomes like depression and life satisfaction, it cannot replace human validation.

Edoardo Sebastiano De Duro, Emma Franchino, Massimo Stella2026-07-27
📄 other

Compartment Mismatch Produces Near-Perfect Apparent Discrimination in Blood-Referenced DNA-Methylation Classifiers: A Public-Cohort Reanalysis of Hematologic-Malignancy Series

This study demonstrates that blood-based DNA-methylation classifiers often achieve near-perfect apparent discrimination in hematologic malignancy cohorts not due to disease detection, but because of unaccounted mismatches between the cellular compartments of the healthy reference and the diseased samples, rendering such uncorrected validations uninterpretable.

Frederic Scheer2026-07-27