Why Does Robustness Reduce Superposition?
Quick summary
arXiv:2608.22155v1 Announce Type: cross Abstract: The study of adversarial examples and their origins remains an open area of research. Mechanistic interpretability, and superposition in particular, offers new avenues for approaching this problem. Gorton & Lewis (2025) demonstrate that adversarial examples arise from superposition and show empirically that adversarial training reduces superposition, yet provide no mechanistic account of why this occurs. We present an empirical explanation inspired by the feature taxonomy of Ilyas et al. (2019), tracing the following chain of causalities: adver
Key takeaways
- arXiv:2608.22155v1 Announce Type: cross Abstract: The study of adversarial examples and their origins remains an open area of research.
- Mechanistic interpretability, and superposition in particular, offers new avenues for approaching this problem.
- Gorton & Lewis (2025) demonstrate that adversarial examples arise from superposition and show empirically that adversarial training reduces superposition, yet provide no mechanistic account of why this occurs.
Why it matters
“Why Does Robustness Reduce Superposition?” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments