arXiv Artificial Intelligence

Calibrating Lightweight Sparse Autoencoder Feature Steering

Calibrating Lightweight Sparse Autoencoder Feature Steering

Quick summary

arXiv:2506.12576v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) can enable inference-time topic steering by modifying latent feature activations, but existing steering methods often fail when target-aligned features are not identified or are modified at the wrong scale. We introduce \textsc{ContrastiveSteer} to address these respective failure modes. First, features are scored by how much more strongly they activate on target-domain text than on general text. Second, steering strength is set using a model-specific calibration. We also introduce \emph{contamination}, a comp

Key takeaways

  • arXiv:2506.12576v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) can enable inference-time topic steering by modifying latent feature activations, but existing steering methods often fail when target-aligned features are not identified or are modified at the wrong scale.
  • We introduce \textsc{ContrastiveSteer} to address these respective failure modes.
  • First, features are scored by how much more strongly they activate on target-domain text than on general text.

Why it matters

This development shows AI moving deeper into everyday software. Productivity potential should be weighed against price, data permissions, exportability and the preservation of human control.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗