Signatures of Steerability in Activation Space of Language Models
Quick summary
arXiv:2609.14151v1 Announce Type: cross Abstract: Steering language models using a set of contrastive representations has been a canonical and computationally efficient method for controlling model behavior. Despite this success in controlling certain model behaviors, the effectiveness of activation steering varies markedly across concepts; the generalization properties of steering vectors are often considered a function of the dataset used to construct them. We make this dataset-dependence claim more rigorous and show that simple separation metrics strongly correlate with the downstream steer
Key takeaways
- arXiv:2609.14151v1 Announce Type: cross Abstract: Steering language models using a set of contrastive representations has been a canonical and computationally efficient method for controlling model behavior.
- Despite this success in controlling certain model behaviors, the effectiveness of activation steering varies markedly across concepts; the generalization properties of steering vectors are often considered a function of the dataset used to construct them.
- We make this dataset-dependence claim more rigorous and show that simple separation metrics strongly correlate with the downstream steer
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments