Linguistics-Aware Non-Distortionary LLM Watermarking
Quick summary
arXiv:2606.00613v2 Announce Type: replace-cross Abstract: Watermarking should identify language-model output without degrading quality or limiting verification to the model provider. Multilingual deployment makes this harder because morphology, segmentation, and script change where watermark evidence can be naturally embedded. We introduce LUNA, a linguistically adaptive watermark that combines model-free detection with single-token non-distortion under the standard random-key model. LUNA estimates normalized next-tag entropy from part-of-speech contexts in an external corpus and uses it to se
Key takeaways
- arXiv:2606.00613v2 Announce Type: replace-cross Abstract: Watermarking should identify language-model output without degrading quality or limiting verification to the model provider.
- Multilingual deployment makes this harder because morphology, segmentation, and script change where watermark evidence can be naturally embedded.
- We introduce LUNA, a linguistically adaptive watermark that combines model-free detection with single-token non-distortion under the standard random-key model.
Why it matters
“Linguistics-Aware Non-Distortionary LLM Watermarking” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments