On BatchNorm Forward Modes in Value-Based Reinforcement Learning
Quick summary
arXiv:2609.06421v1 Announce Type: cross Abstract: Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies report performance degradation in discrete-action value learning on Atari. These failures are surprising because discrete Q-networks lack the action-input distribution mismatch identified by CrossQ. We show for target-based C51 and target-free PQN that the simple choice between running and batch statistics at specific forward passes can reverse this degradation. In C51, switching the BN bootstrap forwar
Key takeaways
- arXiv:2609.06421v1 Announce Type: cross Abstract: Batch normalization (BN) substantially improves sample efficiency in continuous-control actor-critic methods such as CrossQ, yet recent studies report performance degradation in discrete-action value learning on Atari.
- These failures are surprising because discrete Q-networks lack the action-input distribution mismatch identified by CrossQ.
- We show for target-based C51 and target-free PQN that the simple choice between running and batch statistics at specific forward passes can reverse this degradation.
Why it matters
The importance of “On BatchNorm Forward Modes in Value-Based Reinforcement Learning” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments