arXiv Artificial Intelligence

Semantic Chunking and the Entropy of Natural Language

Semantic Chunking and the Entropy of Natural Language

Quick summary

arXiv:2602.13194v3 Announce Type: replace-cross Abstract: Humans and large language models can predict next letter or word from its prior context much better than random guessing, indicating strong redundancy of language viewed as a stochastic process. Quantitatively this redundancy was estimated by Shannon to be around 80\%, which means that every letter of a printed English text conveys approximately 1 bit of information and not 4.8 bits that 27 letters (including spaces) could potentially carry. This estimate was later confirmed by using autoregressive token probabilies computed by large la

Key takeaways

  • arXiv:2602.13194v3 Announce Type: replace-cross Abstract: Humans and large language models can predict next letter or word from its prior context much better than random guessing, indicating strong redundancy of language viewed as a stochastic process.
  • Quantitatively this redundancy was estimated by Shannon to be around 80\%, which means that every letter of a printed English text conveys approximately 1 bit of information and not 4.8 bits that 27 letters (including spaces) could potentially carry.
  • This estimate was later confirmed by using autoregressive token probabilies computed by large la

Why it matters

“Semantic Chunking and the Entropy of Natural Language” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗