arXiv Artificial Intelligence

Softmax Reparameterization for Output-Head Quantization

Softmax Reparameterization for Output-Head Quantization

Quick summary

arXiv:2609.31291v1 Announce Type: cross Abstract: Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ. This one-dimensional search includes the original head and fixed mean-centering, preserves the full-precision softmax distr

Key takeaways

  • arXiv:2609.31291v1 Announce Type: cross Abstract: Large vocabularies make output heads a substantial inference cost in small language models.
  • We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization.
  • The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ.

Why it matters

The importance of “Softmax Reparameterization for Output-Head Quantization” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗