arXiv Artificial Intelligence

Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

Quick summary

arXiv:2609.22512v1 Announce Type: new Abstract: Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and evaluated in similar ways, so they can make the same mistakes. We study how this dependency affects the reliability of consensus. We find substantial error correlation across both open-weight and frontier LLM judges. In our main bank of ten judges, the average pairwise correlation between judge errors is 0.21. As a result, the ten judges only provide roughl

Key takeaways

  • arXiv:2609.22512v1 Announce Type: new Abstract: Consensus among LLM judges is often taken as strong evidence that a decision is correct.
  • This assumes that judges make their errors independently.
  • In practice, LLM judges are often trained and evaluated in similar ways, so they can make the same mistakes.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗