arXiv Artificial Intelligence

Same Answer, Different Confidence: Protocol Sensitivity in LLM Confidence Calibration

Same Answer, Different Confidence: Protocol Sensitivity in LLM Confidence Calibration

Quick summary

arXiv:2605.27752v3 Announce Type: replace Abstract: Is verbalized confidence better calibrated than token likelihood? The answer depends on how the token likelihood is measured: which answer is scored, and under which prompt. Published comparisons diverge on this, and in a twelve-study audit five never state the choice. We fix one prediction event per question, the model's own answer together with its correctness label, and score that same answer under a plain query and inside the confidence prompt, holding the answer and its label fixed. Across four QA datasets and three 7-8B Instruct models

Key takeaways

  • arXiv:2605.27752v3 Announce Type: replace Abstract: Is verbalized confidence better calibrated than token likelihood?
  • The answer depends on how the token likelihood is measured: which answer is scored, and under which prompt.
  • Published comparisons diverge on this, and in a twelve-study audit five never state the choice.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗