GradLev: Token-Parallel Test-Time Training Via Costate Prediction
Quick summary
arXiv:2609.34174v2 Announce Type: replace-cross Abstract: Test-time training (TTT) allows a model to improve its predictions at inference time by updating weights after every observed token. However, sequential gra- dient writes make parallel training difficult. We observe that, given layer inputs and activation gradients (costates), online gradient descent admits exact parallel scans for both forward evaluation and reverse backpropagation. GradLev lever- ages this duality: a causal auxiliary network predicts costates across all tokens in parallel; associative scans compute the adapted weights
Key takeaways
- arXiv:2609.34174v2 Announce Type: replace-cross Abstract: Test-time training (TTT) allows a model to improve its predictions at inference time by updating weights after every observed token.
- However, sequential gra- dient writes make parallel training difficult.
- We observe that, given layer inputs and activation gradients (costates), online gradient descent admits exact parallel scans for both forward evaluation and reverse backpropagation.
Why it matters
“GradLev: Token-Parallel Test-Time Training Via Costate Prediction” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments