Language model agents show in-group trust bias invisible to standard behavioural audits
This study reveals that widely used language model agents exhibit a robust in-group trust bias based on arbitrary labels—a subtle social dynamic that remains undetectable by standard aggregate behavioral audits because it manifests in *who* receives actions rather than *which* actions are chosen.