Solving Without Stopping: On-Policy Distillation at Small Scale
Quick summary
arXiv:2609.37326v1 Announce Type: new Abstract: On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then end the reasoning and answer) and, for comparison, in non-thinking mode (no separate reasoning phase). Long reasoning needs two abilities, solving a problem and knowing when it is solved, and we find that distillation transfers the first, but in thinki
Key takeaways
- arXiv:2609.37326v1 Announce Type: new Abstract: On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models.
- We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then end the reasoning and answer) and, for comparison, in non-thinking mode (no separate reasoning phase).
- Long reasoning needs two abilities, solving a problem and knowing when it is solved, and we find that distillation transfers the first, but in thinki
Why it matters
“Solving Without Stopping: On-Policy Distillation at Small Scale” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Member comments