arXiv Artificial Intelligence

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Quick summary

arXiv:2605.05040v2 Announce Type: replace-cross Abstract: On-policy distillation is an efficient alternative to reinforcement learning, offering dense token-level training signals. However, its reliance on a stronger external teacher has driven recent work on on-policy self-distillation, where the same model serves as both teacher and student under different prompt contexts. Yet, existing self-distillation methods largely reduce learning to KL matching toward the context-augmented teacher model. This approach often suffers from training instability and can degrade reasoning performance over ti

Key takeaways

  • arXiv:2605.05040v2 Announce Type: replace-cross Abstract: On-policy distillation is an efficient alternative to reinforcement learning, offering dense token-level training signals.
  • However, its reliance on a stronger external teacher has driven recent work on on-policy self-distillation, where the same model serves as both teacher and student under different prompt contexts.
  • Yet, existing self-distillation methods largely reduce learning to KL matching toward the context-augmented teacher model.

Why it matters

“Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗