arXiv Artificial Intelligence

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

Quick summary

arXiv:2608.18091v2 Announce Type: replace-cross Abstract: As LLM-as-a-judge becomes increasingly widespread, self-preference -- the tendency of a judge to favor its own outputs -- raises growing concerns about evaluation reliability. However, this bias has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this limitation by changing the object of evaluation: instead of judging generated text, ten LLMs assess sets of narr

Key takeaways

  • arXiv:2608.18091v2 Announce Type: replace-cross Abstract: As LLM-as-a-judge becomes increasingly widespread, self-preference -- the tendency of a judge to favor its own outputs -- raises growing concerns about evaluation reliability.
  • However, this bias has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated.
  • As a result, existing measurements cannot separate genuine self-preference from these confounds.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗