The Price of Token Boundaries: Compression Certificates and Prediction
Quick summary
arXiv:2609.35869v1 Announce Type: new Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported value
Key takeaways
- arXiv:2609.35869v1 Announce Type: new Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries.
- We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule.
- Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported value
Why it matters
The importance of “The Price of Token Boundaries: Compression Certificates and Prediction” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments