arXiv Artificial Intelligence

STEPS: Selective On-Policy Self-Distillation for Reasoning

STEPS: Selective On-Policy Self-Distillation for Reasoning

Quick summary

arXiv:2605.10194v2 Announce Type: replace Abstract: On-policy self-distillation (OPSD) uses a model as its own teacher under privileged context, providing token-level supervision on the model's own reasoning trajectories. However, persistent all-token guidance can degrade training in our math RL setting. Such guidance may unnecessarily constrain non-critical tokens and reinforce biases induced by privileged information unavailable at inference. Motivated by these risks, we propose STEPS, a selective OPSD framework that controls the location, direction, and duration of distillation. STEPS ident

Key takeaways

  • arXiv:2605.10194v2 Announce Type: replace Abstract: On-policy self-distillation (OPSD) uses a model as its own teacher under privileged context, providing token-level supervision on the model's own reasoning trajectories.
  • However, persistent all-token guidance can degrade training in our math RL setting.
  • Such guidance may unnecessarily constrain non-critical tokens and reinforce biases induced by privileged information unavailable at inference.

Why it matters

“STEPS: Selective On-Policy Self-Distillation for Reasoning” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗