Mode-Dependent Rectification for Stable PPO Training
Quick summary
arXiv:2602.05619v2 Announce Type: replace-cross Abstract: Mode-dependent architectural components (layers that behave differently during training and evaluation, such as Batch Normalization or dropout) are commonly used in visual reinforcement learning but can destabilize on-policy optimization. We show that in Proximal Policy Optimization (PPO), discrepancies between training and evaluation behavior induced by Batch Normalization lead to policy mismatch, distributional drift, and reward collapse. We propose Mode-Dependent Rectification (MDR), a lightweight dual-phase training procedure that s
Key takeaways
- arXiv:2602.05619v2 Announce Type: replace-cross Abstract: Mode-dependent architectural components (layers that behave differently during training and evaluation, such as Batch Normalization or dropout) are commonly used in visual reinforcement learning but can destabilize on-policy optimization.
- We show that in Proximal Policy Optimization (PPO), discrepancies between training and evaluation behavior induced by Batch Normalization lead to policy mismatch, distributional drift, and reward collapse.
- We propose Mode-Dependent Rectification (MDR), a lightweight dual-phase training procedure that s
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “Mode-Dependent Rectification for Stable PPO Training” may reshape data collection, model training, output accountability and market access.

Member comments