Suggestible Judges Asymmetric Conformity in Large Language Model Adjudication
This study demonstrates that certain large language models used as adjudicators exhibit asymmetric, presentation-driven conformity to specific label cues (particularly "FALSE" values) rather than impartially evaluating merits, with susceptibility varying significantly across vendors and being mitigated by instruction tuning.