arXiv Artificial Intelligence

Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update

Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update

Quick summary

arXiv:2607.11505v2 Announce Type: replace-cross Abstract: Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward behaviors from policy exploration. While on-policy distillation alleviates this by consolidating independently optimized experts, its reliance on matching absolute expert distributions can yield suboptimal supervision, especially when the target model possesses a different prior or already surpasses the expert's capabilities. To alleviate this, we introduce Proxy OPD (P-OPD), an asynchronous post-train

Key takeaways

  • arXiv:2607.11505v2 Announce Type: replace-cross Abstract: Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward behaviors from policy exploration.
  • While on-policy distillation alleviates this by consolidating independently optimized experts, its reliance on matching absolute expert distributions can yield suboptimal supervision, especially when the target model possesses a different prior or already surpasses the expert's capabilities.
  • To alleviate this, we introduce Proxy OPD (P-OPD), an asynchronous post-train

Why it matters

The significance is not only the legal text but how it changes product design. Decisions around “Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update” may reshape data collection, model training, output accountability and market access.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗