Same Answer, Different Confidence: Protocol Sensitivity in LLM Confidence Calibration
Quick summary
arXiv:2605.27752v3 Announce Type: replace Abstract: Is verbalized confidence better calibrated than token likelihood? The answer depends on how the token likelihood is measured: which answer is scored, and under which prompt. Published comparisons diverge on this, and in a twelve-study audit five never state the choice. We fix one prediction event per question, the model's own answer together with its correctness label, and score that same answer under a plain query and inside the confidence prompt, holding the answer and its label fixed. Across four QA datasets and three 7-8B Instruct models
Key takeaways
- arXiv:2605.27752v3 Announce Type: replace Abstract: Is verbalized confidence better calibrated than token likelihood?
- The answer depends on how the token likelihood is measured: which answer is scored, and under which prompt.
- Published comparisons diverge on this, and in a twelve-study audit five never state the choice.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments