The Surprising Effectiveness of Shared Memory in Looped Transformers
Quick summary
arXiv:2610.02383v1 Announce Type: cross Abstract: Looped Transformers apply the same layers several times per token, adding compute to improve quality without more parameters. Each recursion, however, writes its own key-value cache, so memory still grows with compute. Inference-time techniques can shrink this cache at a cost in quality. We pretrain looped language models to share memory: only the first recursion writes a cache, and later recursions read it while keeping a short window of their own. Surprisingly, we find that sharing memory does not cost quality and instead improves it. At 150M
Key takeaways
- arXiv:2610.02383v1 Announce Type: cross Abstract: Looped Transformers apply the same layers several times per token, adding compute to improve quality without more parameters.
- Each recursion, however, writes its own key-value cache, so memory still grows with compute.
- Inference-time techniques can shrink this cache at a cost in quality.
Why it matters
The importance of “The Surprising Effectiveness of Shared Memory in Looped Transformers” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments