arXiv Artificial Intelligence

Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains

Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains

Quick summary

arXiv:2609.14969v1 Announce Type: cross Abstract: Realignment is a promising approach for improving the cross-lingual transfer ability of multilingual language models, particularly for extremely low-resource languages (LRLs). However, existing realignment methods rely on uniform and random sampling of parallel sentences across languages, which may be suboptimal under limited batch sizes. In practice, models may benefit from seeing certain languages more frequently, especially those that are poorly aligned, and the optimal distribution can evolve throughout training. In this work, we propose a

Key takeaways

  • arXiv:2609.14969v1 Announce Type: cross Abstract: Realignment is a promising approach for improving the cross-lingual transfer ability of multilingual language models, particularly for extremely low-resource languages (LRLs).
  • However, existing realignment methods rely on uniform and random sampling of parallel sentences across languages, which may be suboptimal under limited batch sizes.
  • In practice, models may benefit from seeing certain languages more frequently, especially those that are poorly aligned, and the optimal distribution can evolve throughout training.

Why it matters

The importance of “Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗