Verifier Exploitation in NLI-Guided Iterative Refinement: A Controlled Empirical Analysis
This paper demonstrates that verifier exploitation in NLI-guided iterative refinement is an architectural vulnerability inherent to single-metric feedback loops rather than an optimization artifact, showing how a training-free, zero-gradient pipeline can systematically degrade faithfulness while inflating NLI scores through content truncation, and proposes a universal, annotation-free detection protocol to identify such structural risks.