Optimal Design for Active Preference Learning with Biased LLM Judges
Quick summary
arXiv:2609.38860v1 Announce Type: cross Abstract: Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference learning reduces this cost by selecting informative comparisons, and LLM judges can provide additional scalable feedback. However, the preferences of the judges may deviate from those of the target human population. Even after calibration on trusted reference data, active acquisition can shift the comparison distribution and expose residual judge bias. We therefore incorporate judge deviations into the
Key takeaways
- arXiv:2609.38860v1 Announce Type: cross Abstract: Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly.
- Active preference learning reduces this cost by selecting informative comparisons, and LLM judges can provide additional scalable feedback.
- However, the preferences of the judges may deviate from those of the target human population.
Why it matters
This is more than a company headline: it shows who controls infrastructure, users and data in the AI value chain. The practical effect will appear in product integration, pricing and delivered capacity.

Member comments