📊 statistics

Bayesian Inference for Multi-Source Binary Data with Selection and Measurement Error

This paper introduces a Bayesian modeling framework that jointly addresses misclassification and selection errors in multi-source binary data, demonstrating through simulations and empirical analysis that while composite error can be robustly identified, accurately decomposing it into specific components requires valid indicators, auxiliary variables, and informative priors.

Santiago Gómez-Echeverry, Arnout van Delden, Ton de Waal, Dimitris Pavlopoulos2026-07-30
📊 statistics

A General Framework for Stability Assessment of Non-Deterministic Prediction Models

This paper proposes a general framework for quantifying the stability of non-deterministic prediction models across metric, categorical, and ordinal outcomes by introducing data-based and model-based stability measures that overcome the limitations of existing methods, such as their inability to handle individual test objects and their sensitivity to data composition.

Thomas Martin Lange, Armin Otto Schmitt, Felix Heinrich2026-07-30
📊 statistics

A Sequential Tour-Based Mode Choice Framework with GPS-based Data: Integrating Machine-Learning-Derived Features into Mixed Logit and MNL Models

This study develops a sequential tour-based mode choice framework for Sydney using GPS data and machine-learning-derived features, demonstrating that Mixed Logit models significantly improve prediction for interdependent tour components and that the framework remains robust to upstream estimation uncertainty, thereby validating the use of simulated inputs for practical travel demand forecasting.

Mostafa Rahimi, Maliheh Tabasi, Abdul Rawoof Pinjari, Taha Hossein Rashidi2026-07-30
📊 statistics

Clinical Risk-Weighted Evaluation of ICU Mortality Prediction Models on MIMIC-IV

This paper proposes the Clinical Risk-Weighted Score (CRWS), a novel metric that decouples the penalties for false negatives and false positives to better reflect clinical priorities, demonstrating through MIMIC-IV analysis that it significantly alters model rankings and reveals greater clinical advantages for top-performing algorithms compared to the traditional F1 score.

Aman Chandra H, Abhijna Shivaprakash, Amrutha Sachin Nagvekar, Spandana Sujay, Nagaraja J2026-07-29
📊 statistics

Recovering Parametric Inference for Skewed, Small-Sample, and Heteroscedastic Data: The Coefficient-of-Variation Screen and the Log-Scale Variance Ratio (ρ)

This paper proposes a practical framework using summary statistics to screen for skewness and heteroscedasticity via the coefficient of variation and log-scale variance ratio, enabling researchers to recover valid, more powerful parametric inference for small-sample clinical data where traditional normality tests fail and standard t-tests are deficient.

William J. Dwyer2026-07-28
📊 statistics

T_root: A Comprehensive Closed-Form Statistic for Independence in Sparse and Heterogeneous Contingency Tables

The paper introduces T_root, a novel closed-form, parameter-free statistic that achieves accurate size calibration and robust power for testing independence in contingency tables across a wide range of sparsity and marginal heterogeneity, effectively unifying and outperforming existing methods like Pearson's chi-square and the Cressie-Read statistic while providing a deterministic solution without the need for permutation or bootstrap.

William J. Dwyer2026-07-28
📊 statistics

The Generalized Moment-Transmuted Cauchy Distribution: A Heavy-Tailed Family with Controllable Finite Moments

This paper introduces the Generalized Moment-Transmuted Cauchy (GMTC) distribution, a flexible heavy-tailed model that overcomes the classical Cauchy distribution's limitation of non-existent moments by incorporating a shape parameter to control tail behavior and ensure finite moments, while providing a comprehensive analysis of its properties, estimation methods, and practical applications.

STHITADHI DAS2026-07-28