Multiple latent orderings better predict language model preferences
Quick summary
arXiv:2609.22170v1 Announce Type: cross Abstract: Language models are frequently employed in settings where they are asked to make value judgments and choices. These observed choices often exhibit intransitivity: A model may prefer item $A$ to $B$ and $B$ to $C$, while also preferring $C$ to $A$. Existing work that models LLM preferences treats such inconsistencies as sampling noise around a single latent ordering. We instead propose that intransitivity reflects the aggregation of multiple latent, internally consistent orderings. We first show that observed inconsistencies cannot be explained
Key takeaways
- arXiv:2609.22170v1 Announce Type: cross Abstract: Language models are frequently employed in settings where they are asked to make value judgments and choices.
- These observed choices often exhibit intransitivity: A model may prefer item $A$ to $B$ and $B$ to $C$, while also preferring $C$ to $A$.
- Existing work that models LLM preferences treats such inconsistencies as sampling noise around a single latent ordering.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments