arXiv Artificial Intelligence

Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving

Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving

Quick summary

arXiv:2609.36750v1 Announce Type: cross Abstract: Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Ad

Key takeaways

  • arXiv:2609.36750v1 Announce Type: cross Abstract: Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels.
  • Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly.
  • However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group.

Why it matters

The significance is not only the legal text but how it changes product design. Decisions around “Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving” may reshape data collection, model training, output accountability and market access.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗