arXiv Artificial Intelligence

Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging

Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging

Quick summary

arXiv:2609.30751v1 Announce Type: new Abstract: Pairwise language-model judges can gather evidence through direct comparison, reasoning, or reference-based verification, but no single protocol is best across benchmarks and judge backbones. We introduce Backbone-Adaptive Evidence Routing (BAER), which adapts the evidence mechanism while preserving candidate symmetry: swapping the two responses may reverse the preference but cannot change its strength. BAER separates each expert's signed preference from candidate-invariant reliability and builds three symmetric heads: evidence stacking, reliabil

Key takeaways

  • arXiv:2609.30751v1 Announce Type: new Abstract: Pairwise language-model judges can gather evidence through direct comparison, reasoning, or reference-based verification, but no single protocol is best across benchmarks and judge backbones.
  • We introduce Backbone-Adaptive Evidence Routing (BAER), which adapts the evidence mechanism while preserving candidate symmetry: swapping the two responses may reverse the preference but cannot change its strength.
  • BAER separates each expert's signed preference from candidate-invariant reliability and builds three symmetric heads: evidence stacking, reliabil

Why it matters

“Backbone-Adaptive Evidence Routing for Robust Pairwise LLM Judging” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗