arXiv Artificial Intelligence

Objective vs. Search: Decomposing What Makes a Good Tokeniser

Objective vs. Search: Decomposing What Makes a Good Tokeniser

Quick summary

arXiv:2609.19145v1 Announce Type: cross Abstract: Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM. These differ along two orthogonal axes: their optimisation objective (compression vs. log-likelihood) and their search procedure (bottom-up merging vs. top-down pruning). Existing comparisons confound these axes, making it unclear whether their observed differences stem from what is being optimised vs. how it is being optimised. We disentangle the two by introducing two new tokenisation algorithms that complete this 2x2 design space:

Key takeaways

  • arXiv:2609.19145v1 Announce Type: cross Abstract: Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM.
  • These differ along two orthogonal axes: their optimisation objective (compression vs.
  • log-likelihood) and their search procedure (bottom-up merging vs.

Why it matters

“Objective vs. Search: Decomposing What Makes a Good Tokeniser” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗