Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers
Quick summary
arXiv:2609.18321v1 Announce Type: cross Abstract: Offline on-policy distillation gains efficiency by collecting student trajectories and teacher supervision once and reusing them throughout optimization. The same reuse makes imperfect supervision persistent. Since even strong teachers can fail, we ask \emph{what remains learnable from imperfect teacher supervision?} Teacher failure is only a coarse problem-level signal and does not imply that all supervision along the associated student trajectory is unhelpful. A natural alternative is to estimate teacher recoverability along the trajectory, b
Key takeaways
- arXiv:2609.18321v1 Announce Type: cross Abstract: Offline on-policy distillation gains efficiency by collecting student trajectories and teacher supervision once and reusing them throughout optimization.
- The same reuse makes imperfect supervision persistent.
- Since even strong teachers can fail, we ask \emph{what remains learnable from imperfect teacher supervision?} Teacher failure is only a coarse problem-level signal and does not imply that all supervision along the associated student trajectory is unhelpful.
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers” may reshape data collection, model training, output accountability and market access.

Member comments