The Curse of Multilinguality in Lexical Normalization
Quick summary
arXiv:2609.00329v1 Announce Type: cross Abstract: Lexical normalization rewrites the noisy, non-standard words that fill user-generated text (tmrw, u, gr8) into their standard forms. Because labelled data is scarce for most languages, a popular shortcut is to train a single model on many languages at once. We ask a simple question: how many languages should such a model be trained on? Using one fixed-capacity character-level model and twelve languages from a standard benchmark, we vary the number of jointly trained languages from one to twelve and measure per-language accuracy. We find a clear
Key takeaways
- arXiv:2609.00329v1 Announce Type: cross Abstract: Lexical normalization rewrites the noisy, non-standard words that fill user-generated text (tmrw, u, gr8) into their standard forms.
- Because labelled data is scarce for most languages, a popular shortcut is to train a single model on many languages at once.
- We ask a simple question: how many languages should such a model be trained on?
Why it matters
“The Curse of Multilinguality in Lexical Normalization” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments