ReCo: Reweighting GRPO Against Distributional Concentration
Quick summary
arXiv:2607.26862v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage of reasoning paths. We find that this reduction is associated with GRPO concentrating on responses that the base model already generates with high probability. We trace this concentration to two mechanisms in the GRPO update. At the response level, high-probability respo
Key takeaways
- arXiv:2607.26862v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models.
- Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage of reasoning paths.
- We find that this reduction is associated with GRPO concentrating on responses that the base model already generates with high probability.
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “ReCo: Reweighting GRPO Against Distributional Concentration” may reshape data collection, model training, output accountability and market access.
