Why Pretraining Fails to Share Cross-Lingual Knowledge
Quick summary
arXiv:2609.19291v1 Announce Type: cross Abstract: Large Language Models (LLMs) have made remarkable progress in the processing and modeling of many languages. Yet, unlike human multilinguals, they exhibit surprisingly limited cross-lingual knowledge transfer. While this limitation is well documented, its origins during multilingual training remain unclear. We pretrain 360M- and 7B-parameter LLMs and show that poor cross-lingual knowledge generalization emerges during pretraining and persists under standard interventions. To isolate its cause, we employ a controlled bilingual pretraining settin
Key takeaways
- arXiv:2609.19291v1 Announce Type: cross Abstract: Large Language Models (LLMs) have made remarkable progress in the processing and modeling of many languages.
- Yet, unlike human multilinguals, they exhibit surprisingly limited cross-lingual knowledge transfer.
- While this limitation is well documented, its origins during multilingual training remain unclear.
Why it matters
“Why Pretraining Fails to Share Cross-Lingual Knowledge” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Member comments