arXiv Artificial Intelligence

Measuring Semantic Abstractness of SAE Features via Nonlocality

Measuring Semantic Abstractness of SAE Features via Nonlocality

Quick summary

arXiv:2608.10537v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature. To this end, we introduce \emph{Feature Nonlocality} (FNL), defined

Key takeaways

  • arXiv:2608.10537v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features.
  • To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones.
  • However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature.

Why it matters

“Measuring Semantic Abstractness of SAE Features via Nonlocality” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗