📊 statistics

Reliable AUC Evaluation for Positive-Unlabeled Classifiers: Calibrated Confidence Intervals under an Unknown Class Prior

This paper proposes a method to derive calibrated, two-sided confidence intervals for the true Area Under the Curve (AUC) in Positive-Unlabeled learning by exactly recovering the target AUC from observable metrics and propagating the uncertainty of the estimated positive fraction, thereby addressing the bias and lack of reliability in current performance evaluations.

Vincent Looten2026-09-01
📊 statistics

Temporal Modeling of Weekly Dengue Cases in Marikina City: A Time-Respecting Comparison of Classical, Machine Learning, and Deep Temporal Models

This study evaluates classical, machine learning, and deep learning models for forecasting weekly dengue cases in Marikina City using a leakage-safe temporal design, finding that while simple baselines performed well during internal validation, the LSTM model achieved superior predictive accuracy on the 2025 out-of-time holdout despite a general tendency to underpredict outbreak-level weeks.

Kal-el John C. Bautista, Luis Arnold N. Respecio, Cyril Troy C. Vinuya, John Paul Q. Tomas, Karl C. Ablola, Bonifacio T. (…)2026-08-31
📊 statistics

Participant-Specific Voice-Presenter Interactions in Short Learning Videos: An Uncertainty-Aware Bayesian Analysis of AI-Generated and Human-Produced Components

This study employs a participant-specific, uncertainty-aware Bayesian hierarchical model to analyze a within-participant experiment, revealing that AI-generated voices paired with AI avatars yield the most favorable engagement and cognitive load profile while highlighting significant individual heterogeneity in responses to audiovisual combinations.

Yanshuai Shi2026-08-31
📊 statistics

Automated auditing of statistical quality in university theses: why lexical rules fail and what a language layer adds

This study demonstrates that rule-based automated auditing of statistical quality in university theses performs worse than chance due to fundamental lexical limitations, but integrating a language model layer significantly improves agreement with expert judgment while maintaining high recall and low cost.

Milton Vladimir Mamani Calisaya, Charles Ignacio Mendoza Mollocondo, Vladimiro Ibañez Quispe, Jesus Pari Flores2026-08-28
📊 statistics

Grade 4-to-5 Joint Mathematics–Science Benchmark Transitions in Nine Systems: A Descriptive Common-Severity Analysis Using TIMSS 2023 Longitudinal Data

This study utilizes TIMSS 2023 longitudinal data from nine systems to estimate and standardize the joint probability of students transitioning from below to above both mathematics and science benchmarks, revealing significant system-specific variations and demonstrating that while common-severity standardization clarifies target dependence, the results remain descriptive rather than causal.

Jonas A. Mandalunes2026-08-28
📊 statistics

When the adjustment fails: population-denominator sensitivity in coverage-adjusted PISA trends

This paper demonstrates that PISA's coverage-adjusted top-quarter trends are highly sensitive to the choice of population denominator, revealing significant discrepancies between reported data and World Population Prospects estimates that underscore the need for transparent provenance reporting and caution when synthetic weights trigger adjustments.

Jonas A. Mandalunes, Danica Jane S. Mandalunes2026-08-28
📊 statistics

Measurement-bounded predictive inference in international large-scale assessment: how the plausible-value ceiling governs attainable accuracy and the stability of predictor rankings in PISA 2022

This study demonstrates that the measurement precision ceiling inherent in PISA 2022 plausible values fundamentally limits attainable predictive accuracy and destabilizes predictor rankings across domains and education systems, necessitating that performance comparisons and model interpretations be grounded in these computable measurement bounds.

Jonas A. Mandalunes2026-08-28