arXiv Artificial Intelligence

Calibrating Teacher--Student Discrepancy for On-Policy Distillation

Calibrating Teacher--Student Discrepancy for On-Policy Distillation

Quick summary

arXiv:2609.21619v1 Announce Type: new Abstract: On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the teacher itself, which are consequently mixed into the observed teacher--student discrepancy and indiscriminately learned by standard OPD during training. This issue is further exacerbated by privileged OPD, where privileged information induces larger te

Key takeaways

  • arXiv:2609.21619v1 Announce Type: new Abstract: On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student.
  • However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the teacher itself, which are consequently mixed into the observed teacher--student discrepancy and indiscriminately learned by standard OPD during training.
  • This issue is further exacerbated by privileged OPD, where privileged information induces larger te

Why it matters

The significance is not only the legal text but how it changes product design. Decisions around “Calibrating Teacher--Student Discrepancy for On-Policy Distillation” may reshape data collection, model training, output accountability and market access.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗