arXiv Artificial Intelligence

ReCo: Reweighting GRPO Against Distributional Concentration

ReCo: Reweighting GRPO Against Distributional Concentration

Quick summary

arXiv:2607.26862v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage of reasoning paths. We find that this reduction is associated with GRPO concentrating on responses that the base model already generates with high probability. We trace this concentration to two mechanisms in the GRPO update. At the response level, high-probability respo

Key takeaways

  • arXiv:2607.26862v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models.
  • Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage of reasoning paths.
  • We find that this reduction is associated with GRPO concentrating on responses that the base model already generates with high probability.

Why it matters

The significance is not only the legal text but how it changes product design. Decisions around “ReCo: Reweighting GRPO Against Distributional Concentration” may reshape data collection, model training, output accountability and market access.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗