An LLM-Native Psychometric Instrument Does Not Predict LLM Behavior: Evidence Across 25 Models
Despite demonstrating high internal reliability, a novel psychometric instrument derived from LLM behavioral affordances fails to predict actual LLM behavior, revealing a critical disconnect between model self-reports and observed actions that poses a significant risk for LLM-as-judge evaluation pipelines.