Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
Quick summary
arXiv:2608.18091v2 Announce Type: replace-cross Abstract: As LLM-as-a-judge becomes increasingly widespread, self-preference -- the tendency of a judge to favor its own outputs -- raises growing concerns about evaluation reliability. However, this bias has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this limitation by changing the object of evaluation: instead of judging generated text, ten LLMs assess sets of narr
Key takeaways
- arXiv:2608.18091v2 Announce Type: replace-cross Abstract: As LLM-as-a-judge becomes increasingly widespread, self-preference -- the tendency of a judge to favor its own outputs -- raises growing concerns about evaluation reliability.
- However, this bias has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated.
- As a result, existing measurements cannot separate genuine self-preference from these confounds.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments