arXiv Artificial Intelligence

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

Quick summary

arXiv:2608.11947v1 Announce Type: cross Abstract: Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge. In this paper, we test whether preventing a model from seeing option labels while committing to an answer removes positional influence and, in turn, improves performance. We evaluate two different strategies for mitigating bias. The first uses a generation-then-matching approach, and the second scores options in isolation, which is positionally

Key takeaways

  • arXiv:2608.11947v1 Announce Type: cross Abstract: Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge.
  • In this paper, we test whether preventing a model from seeing option labels while committing to an answer removes positional influence and, in turn, improves performance.
  • We evaluate two different strategies for mitigating bias.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗