arXiv Artificial Intelligence

JudgeProfile: Understanding and Steering Subjectivity in LLM Judges

JudgeProfile: Understanding and Steering Subjectivity in LLM Judges

Quick summary

arXiv:2609.36705v1 Announce Type: new Abstract: LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specific attributes like clarity, correctness, and detail) and prioritization (how much each attribute influences the final choice). We curate SubjectiveSet, a dataset of 50,013 response pairs from 17 public data sources, evaluated by 21 LLM judges across 87

Key takeaways

  • arXiv:2609.36705v1 Announce Type: new Abstract: LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong.
  • To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specific attributes like clarity, correctness, and detail) and prioritization (how much each attribute influences the final choice).
  • We curate SubjectiveSet, a dataset of 50,013 response pairs from 17 public data sources, evaluated by 21 LLM judges across 87

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗