Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update
Quick summary
arXiv:2607.11505v2 Announce Type: replace-cross Abstract: Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward behaviors from policy exploration. While on-policy distillation alleviates this by consolidating independently optimized experts, its reliance on matching absolute expert distributions can yield suboptimal supervision, especially when the target model possesses a different prior or already surpasses the expert's capabilities. To alleviate this, we introduce Proxy OPD (P-OPD), an asynchronous post-train
Key takeaways
- arXiv:2607.11505v2 Announce Type: replace-cross Abstract: Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward behaviors from policy exploration.
- While on-policy distillation alleviates this by consolidating independently optimized experts, its reliance on matching absolute expert distributions can yield suboptimal supervision, especially when the target model possesses a different prior or already surpasses the expert's capabilities.
- To alleviate this, we introduce Proxy OPD (P-OPD), an asynchronous post-train
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update” may reshape data collection, model training, output accountability and market access.

Member comments