Output Embedding Centering for Stable LLM Pretraining
Quick summary
arXiv:2601.02031v3 Announce Type: replace-cross Abstract: Pretraining of large language models is not only expensive but also prone to certain training instabilities. A specific instability that often occurs at the end of training is output logit divergence. The most widely used mitigation strategies, z-loss and logit soft-capping, merely address the symptoms rather than the underlying cause of the problem. In this paper, we analyze the instability from the perspective of the output embeddings' geometry and identify anisotropic embeddings as its source. Based on this, we propose output embeddi
Key takeaways
- arXiv:2601.02031v3 Announce Type: replace-cross Abstract: Pretraining of large language models is not only expensive but also prone to certain training instabilities.
- A specific instability that often occurs at the end of training is output logit divergence.
- The most widely used mitigation strategies, z-loss and logit soft-capping, merely address the symptoms rather than the underlying cause of the problem.
Why it matters
“Output Embedding Centering for Stable LLM Pretraining” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments