arXiv Artificial Intelligence

Rescaling Confidence: What Scale Design Reveals About LLM Metacognition

Rescaling Confidence: What Scale Design Reveals About LLM Metacognition

Quick summary

arXiv:2603.09309v3 Announce Type: replace Abstract: Verbalized confidence, in which LLMs report a numerical certainty score, is widely used to estimate uncertainty in black-box settings, yet the confidence scale itself (typically 0--100) is rarely examined. We show that this design choice is not neutral. Across six LLMs and three datasets, verbalized confidence is heavily discretized, with more than 78\% of responses concentrating on just three round-number values. To investigate this phenomenon, we systematically manipulate confidence scales along three dimensions: granularity, boundary place

Key takeaways

  • arXiv:2603.09309v3 Announce Type: replace Abstract: Verbalized confidence, in which LLMs report a numerical certainty score, is widely used to estimate uncertainty in black-box settings, yet the confidence scale itself (typically 0--100) is rarely examined.
  • We show that this design choice is not neutral.
  • Across six LLMs and three datasets, verbalized confidence is heavily discretized, with more than 78\% of responses concentrating on just three round-number values.

Why it matters

The importance of “Rescaling Confidence: What Scale Design Reveals About LLM Metacognition” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗