arXiv Artificial Intelligence

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models

Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models

Quick summary

arXiv:2607.18292v3 Announce Type: replace-cross Abstract: Bigger language models are less reliable. Across three families, three benchmarks and six rungs, including in-the-wild chat logs, scaling closes the start-of-response knowledge gap up to $7\times$ while within-response knowledge degradation grows up to $39\times$. We trace that residual to one variable, the per-position disagreement $\delta = \log p_M - \log p_O$ against a stronger oracle, whose second moment splits exactly into bias$^2$ $\mathrm{KL}(p_M \,\|\, p_O)^2$ and decoding risk $\mathrm{Var}[\delta]$. That split is an interpret

Key takeaways

  • arXiv:2607.18292v3 Announce Type: replace-cross Abstract: Bigger language models are less reliable.
  • Across three families, three benchmarks and six rungs, including in-the-wild chat logs, scaling closes the start-of-response knowledge gap up to $7\times$ while within-response knowledge degradation grows up to $39\times$.
  • We trace that residual to one variable, the per-position disagreement $\delta = \log p_M - \log p_O$ against a stronger oracle, whose second moment splits exactly into bias$^2$ $\mathrm{KL}(p_M \,\|\, p_O)^2$ and decoding risk $\mathrm{Var}[\delta]$.

Why it matters

The importance of “Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗