arXiv Artificial Intelligence

The Curse of Multilinguality in Lexical Normalization

The Curse of Multilinguality in Lexical Normalization

Quick summary

arXiv:2609.00329v1 Announce Type: cross Abstract: Lexical normalization rewrites the noisy, non-standard words that fill user-generated text (tmrw, u, gr8) into their standard forms. Because labelled data is scarce for most languages, a popular shortcut is to train a single model on many languages at once. We ask a simple question: how many languages should such a model be trained on? Using one fixed-capacity character-level model and twelve languages from a standard benchmark, we vary the number of jointly trained languages from one to twelve and measure per-language accuracy. We find a clear

Key takeaways

  • arXiv:2609.00329v1 Announce Type: cross Abstract: Lexical normalization rewrites the noisy, non-standard words that fill user-generated text (tmrw, u, gr8) into their standard forms.
  • Because labelled data is scarce for most languages, a popular shortcut is to train a single model on many languages at once.
  • We ask a simple question: how many languages should such a model be trained on?

Why it matters

“The Curse of Multilinguality in Lexical Normalization” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗