arXiv Artificial Intelligence

Signatures of Steerability in Activation Space of Language Models

Signatures of Steerability in Activation Space of Language Models

Quick summary

arXiv:2609.14151v1 Announce Type: cross Abstract: Steering language models using a set of contrastive representations has been a canonical and computationally efficient method for controlling model behavior. Despite this success in controlling certain model behaviors, the effectiveness of activation steering varies markedly across concepts; the generalization properties of steering vectors are often considered a function of the dataset used to construct them. We make this dataset-dependence claim more rigorous and show that simple separation metrics strongly correlate with the downstream steer

Key takeaways

  • arXiv:2609.14151v1 Announce Type: cross Abstract: Steering language models using a set of contrastive representations has been a canonical and computationally efficient method for controlling model behavior.
  • Despite this success in controlling certain model behaviors, the effectiveness of activation steering varies markedly across concepts; the generalization properties of steering vectors are often considered a function of the dataset used to construct them.
  • We make this dataset-dependence claim more rigorous and show that simple separation metrics strongly correlate with the downstream steer

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗