$S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient
Quick summary
arXiv:2609.37976v1 Announce Type: cross Abstract: LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training. Unlike existing efforts that mostly operate within the dominant subspace, we are the
Key takeaways
- arXiv:2609.37976v1 Announce Type: cross Abstract: LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost.
- We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training.
- Unlike existing efforts that mostly operate within the dominant subspace, we are the
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments