The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean
This paper identifies a "double measurement confound" in LLM agent benchmarks where fixed scaffolds and flawed scorers obscure true model capability, and proposes a measurement-theoretic framework with an audit-and-repair protocol to transfer execution control to models, implement ground-truth scoring, and report reliability metrics beyond the mean.