Finite-Horizon Input-Output Dynamics of Minibatch Perturbations in AdamW
Quick summary
arXiv:2608.19762v1 Announce Type: cross Abstract: A minibatch can influence training beyond the update at which it is observed because AdamW stores past gradient information in its optimizer states. We study this delayed effect through paired trajectories that differ only in one gradient update and share the same subsequent training sequence. We formulate AdamW as a finite-horizon input--state--output (ISO) system whose state contains the model parameters and first- and second-moment estimates. Linearizing the joint dynamics yields a signed response operator that maps a localized gradient pert
Key takeaways
- arXiv:2608.19762v1 Announce Type: cross Abstract: A minibatch can influence training beyond the update at which it is observed because AdamW stores past gradient information in its optimizer states.
- We study this delayed effect through paired trajectories that differ only in one gradient update and share the same subsequent training sequence.
- We formulate AdamW as a finite-horizon input--state--output (ISO) system whose state contains the model parameters and first- and second-moment estimates.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments