Calibrating Teacher--Student Discrepancy for On-Policy Distillation
Quick summary
arXiv:2609.21619v1 Announce Type: new Abstract: On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the teacher itself, which are consequently mixed into the observed teacher--student discrepancy and indiscriminately learned by standard OPD during training. This issue is further exacerbated by privileged OPD, where privileged information induces larger te
Key takeaways
- arXiv:2609.21619v1 Announce Type: new Abstract: On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student.
- However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the teacher itself, which are consequently mixed into the observed teacher--student discrepancy and indiscriminately learned by standard OPD during training.
- This issue is further exacerbated by privileged OPD, where privileged information induces larger te
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “Calibrating Teacher--Student Discrepancy for On-Policy Distillation” may reshape data collection, model training, output accountability and market access.

Member comments