SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation
Quick summary
arXiv:2609.36742v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitiga
Key takeaways
- arXiv:2609.36742v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps.
- To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals.
- However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice.
Why it matters
“SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Member comments