arXiv Artificial Intelligence

Minimally Invasive Steering of Language Models

Minimally Invasive Steering of Language Models

Quick summary

arXiv:2609.30218v1 Announce Type: cross Abstract: Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states. Unregularized reward optimization can substantially alter the output distribution and degrade generation quality. We propose Minimally Invasive Steering Vector Optimization (MISVO), which penalizes interventions using the local KL geometry of the induced token distribution. The resulting Fisher quadratic measures distributional sensitivity and admits an analytic gradient computed through matrix--vector products with the frozen l

Key takeaways

  • arXiv:2609.30218v1 Announce Type: cross Abstract: Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states.
  • Unregularized reward optimization can substantially alter the output distribution and degrade generation quality.
  • We propose Minimally Invasive Steering Vector Optimization (MISVO), which penalizes interventions using the local KL geometry of the induced token distribution.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗