Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models
Quick summary
arXiv:2607.18292v3 Announce Type: replace-cross Abstract: Bigger language models are less reliable. Across three families, three benchmarks and six rungs, including in-the-wild chat logs, scaling closes the start-of-response knowledge gap up to $7\times$ while within-response knowledge degradation grows up to $39\times$. We trace that residual to one variable, the per-position disagreement $\delta = \log p_M - \log p_O$ against a stronger oracle, whose second moment splits exactly into bias$^2$ $\mathrm{KL}(p_M \,\|\, p_O)^2$ and decoding risk $\mathrm{Var}[\delta]$. That split is an interpret
Key takeaways
- arXiv:2607.18292v3 Announce Type: replace-cross Abstract: Bigger language models are less reliable.
- Across three families, three benchmarks and six rungs, including in-the-wild chat logs, scaling closes the start-of-response knowledge gap up to $7\times$ while within-response knowledge degradation grows up to $39\times$.
- We trace that residual to one variable, the per-position disagreement $\delta = \log p_M - \log p_O$ against a stronger oracle, whose second moment splits exactly into bias$^2$ $\mathrm{KL}(p_M \,\|\, p_O)^2$ and decoding risk $\mathrm{Var}[\delta]$.
Why it matters
The importance of “Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.
