arXiv Artificial Intelligence

The Price of Token Boundaries: Compression Certificates and Prediction

The Price of Token Boundaries: Compression Certificates and Prediction

Quick summary

arXiv:2609.35869v1 Announce Type: new Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported value

Key takeaways

  • arXiv:2609.35869v1 Announce Type: new Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries.
  • We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule.
  • Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported value

Why it matters

The importance of “The Price of Token Boundaries: Compression Certificates and Prediction” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗