arXiv Artificial Intelligence

On the Complexity of Preference-Based Bandits

On the Complexity of Preference-Based Bandits

Quick summary

arXiv:2609.39351v1 Announce Type: new Abstract: We study preference-based bandits with general reward function classes, where a learner sequentially selects pairs of arms and observes binary preference feedback governed by the Bradley--Terry model. This setting naturally arises in applications such as recommender systems, tournament ranking, and learning from human feedback, where relative preferences are easier to elicit than absolute rewards. The observation model inherits the logistic bandit challenge of handling the problem-dependent constant $\kappa$, which accounts for the non-linearity

Key takeaways

  • arXiv:2609.39351v1 Announce Type: new Abstract: We study preference-based bandits with general reward function classes, where a learner sequentially selects pairs of arms and observes binary preference feedback governed by the Bradley--Terry model.
  • This setting naturally arises in applications such as recommender systems, tournament ranking, and learning from human feedback, where relative preferences are easier to elicit than absolute rewards.
  • The observation model inherits the logistic bandit challenge of handling the problem-dependent constant $\kappa$, which accounts for the non-linearity

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗