Recurrent Looped Transformer
Quick summary
arXiv:2610.07591v1 Announce Type: cross Abstract: State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length. We introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encoder output with the previous token's final decoder state, so the computation path grows with sequence length at a fixed per-token cost. On six algorithmic tasks, we compare five splits of eight layers with an eight-layer Transformer
Key takeaways
- arXiv:2610.07591v1 Announce Type: cross Abstract: State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length.
- We introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder.
- At each token, the decoder merges the encoder output with the previous token's final decoder state, so the computation path grows with sequence length at a fixed per-token cost.
Why it matters
“Recurrent Looped Transformer” is a product decision that may change how people work with AI. Its value depends on task completion, correction effort and data handling—not simply the presence of a new feature.

Member comments