arXiv Artificial Intelligence

Stability of Transformers under Layer Normalization

Stability of Transformers under Layer Normalization

Quick summary

arXiv:2510.09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable. Layer normalization, a standard component, improves training stability, but its placement has often been ad-hoc. In this paper, we conduct a principled study on the forward (hidden states) and backward (gradient) stability of Transformers under different layer normalization placements. Our theory provides key insights into the training dynamics: whether training drives Transformers toward regular solutions or pathological behaviors. For forward stability, we deriv

Key takeaways

  • arXiv:2510.09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable.
  • Layer normalization, a standard component, improves training stability, but its placement has often been ad-hoc.
  • In this paper, we conduct a principled study on the forward (hidden states) and backward (gradient) stability of Transformers under different layer normalization placements.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗