Widespread use of invalid statistical tests in biomedical machine learning
This paper reveals that the widespread use of invalid statistical tests ignoring cross-validation fold dependence in biomedical machine learning leads to inflated false positive rates, prompting the authors to propose the SHARP test as a robust solution and provide new reporting guidelines for valid model comparison.