Detector-Calibration Failures in Pattern-Based LLM Refusal Classification: Discovery, Generalization, and a Confirmed False-Positive Pattern Across Models
This paper details a three-phase investigation revealing that apparent non-determinism in LLM refusal detection was largely caused by correctable detector artifacts, which, while generalizable across models, introduce specific false-positive patterns that necessitate manual auditing and transparent reporting of data gaps to ensure accurate safety assessments.