Data-Free Pruning of Self-Attention Layers in LLMs
Quick summary
arXiv:2512.20636v2 Announce Type: replace-cross Abstract: Many self-attention sublayers in large language models (LLMs) can be removed with little to no loss. We attribute this to the Attention Suppression Hypothesis: during pre-training, some deep attention layers learn to mute their own contribution, leaving the residual stream and the MLP to carry the representation. We propose Gate-Norm, a one-shot, weight-only criterion that ranks attention sublayers by query-key coupling and removes the least coupled ones, requiring no calibration data, no forward passes, no fine-tuning, and no specializ
Key takeaways
- arXiv:2512.20636v2 Announce Type: replace-cross Abstract: Many self-attention sublayers in large language models (LLMs) can be removed with little to no loss.
- We attribute this to the Attention Suppression Hypothesis: during pre-training, some deep attention layers learn to mute their own contribution, leaving the residual stream and the MLP to carry the representation.
- We propose Gate-Norm, a one-shot, weight-only criterion that ranks attention sublayers by query-key coupling and removes the least coupled ones, requiring no calibration data, no forward passes, no fine-tuning, and no specializ
Why it matters
The importance of “Data-Free Pruning of Self-Attention Layers in LLMs” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments