arXiv Artificial Intelligence

Cost-Aware Best-LLM Identification using Dueling Feedback

Cost-Aware Best-LLM Identification using Dueling Feedback

Quick summary

arXiv:2609.30360v1 Announce Type: cross Abstract: Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons between model responses provide robust preference signals, and (ii) heterogeneous sampling costs, reflecting the differing costs of querying different LLMs. Assuming the existence of a Condorcet winner, a condition we empirically validate across multiple real-world datasets, we propose a T

Key takeaways

  • arXiv:2609.30360v1 Announce Type: cross Abstract: Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons between model responses provide robust preference signals, and (ii) heterogeneous sampling costs, reflecting the differing costs of querying different LLMs.
  • Assuming the existence of a Condorcet winner, a condition we empirically validate across multiple real-world datasets, we propose a T

Why it matters

“Cost-Aware Best-LLM Identification using Dueling Feedback” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗