Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
Quick summary
arXiv:2609.18708v1 Announce Type: cross Abstract: In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Carlo continuations, change sharply across intermediate states while critic predictions remain comparatively flat. We further observe this phenomenon in a controlled FrozenLake environment and find that it becomes more pronounced as
Key takeaways
- arXiv:2609.18708v1 Announce Type: cross Abstract: In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates.
- However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Carlo continuations, change sharply across intermediate states while critic predictions remain comparatively flat.
- We further observe this phenomenon in a controlled FrozenLake environment and find that it becomes more pronounced as
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening” may reshape data collection, model training, output accountability and market access.

Member comments