This collection explores the rapidly evolving intersection of statistics and artificial intelligence, where mathematical rigor meets machine learning innovation. These studies tackle how we can trust, interpret, and improve the algorithms that power modern technology, moving beyond black-box models to create systems that are both powerful and understandable. By bridging theoretical statistics with practical AI applications, researchers here are defining the next generation of intelligent tools.

Gist.Science processes every new preprint in this category directly from arXiv, ensuring you have immediate access to the latest breakthroughs before they appear in formal journals. For each paper, we provide both a detailed technical breakdown for experts and a clear, plain-language summary for anyone curious about the future of data and intelligence. Below are the latest papers in statistics and artificial intelligence, curated to help you stay ahead of the curve.

📊 statistics

Decoupling risk and masking in mammographic density under irregular follow up using a latent Markov progression detection framework

This paper proposes a two-phase latent Markov framework that decouples breast cancer risk from mammographic masking by modeling an underlying latent disease progression separate from observed density, thereby quantifying how irregular screening and masking effects can significantly underestimate risk in specific patient groups.

Furkan Danisman, Zarina Oflaz, Zeynep Kalaylioglu, Mahmut Onur Kulturoglu, Lutfi Dogan2026-07-21
⚡ electrical engineering

Approximate Relative Entropy Constraints for Nonlinear Covariance Steering Under Distribution Ambiguity

This paper proposes a distributionally robust covariance-steering framework that utilizes relative entropy constraints and computable upper bounds on risk-sensitive quantities to design stochastic guidance policies for nonlinear systems, ensuring the true state distribution remains close to a Gaussian surrogate while mitigating estimation errors caused by distribution ambiguity.

Trevor N. Wolf, Jay W. McMahon2026-07-21
📊 statistics

National Versus Domain: Coverage Properties of HB Credible Intervals Under Survey Redesign

This paper demonstrates that while hierarchical Bayes (HB) credible intervals maintain near-nominal national coverage and significantly outperform classical direct estimators under stress-test scenarios involving unsampled strata, their domain-level coverage for Employment and Unemployment remains below nominal due to shrinkage bias that cannot be corrected by prior sensitivity or MSE adjustments.

Siu-Ming Tam2026-07-21
📊 statistics

Comparing Missing Data Methods for Estimating Average Treatment Effects Under Time-Varying Confounding: A Simulation Study

This simulation study evaluates the performance of various missing data imputation methods for estimating average treatment effects under time-varying confounding, finding that multiple imputation generally outperforms other approaches in reducing bias and improving coverage, particularly when missingness mechanisms are complex.

Ben Swallow, Lars Brestrich, Victor Velasco-Pardo2026-07-21
📊 statistics

PIONEER: Bayesian Joint Modelling of Mechanistic Tumour Growth and Time-to-Event Endpoints for Dynamic Prediction of Ongoing Oncology Trials

The paper introduces PIONEER, a Bayesian joint modelling framework that integrates mechanistic tumour growth dynamics with multistate survival analysis to enable calibrated, uncertainty-quantified forecasting of clinical trial endpoints like PFS and OS using immature data and early longitudinal tumour measurements.

Karim Naguib, Roger Berché, Lu Li, Antonia Bevan, Sajan Khosla, Jessica Davies, Paul Metcalfe2026-07-21