When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
Quick summary
arXiv:2608.03632v1 Announce Type: new Abstract: On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable. However, the assumptions overlook a fundamental failure mode of language models: their token-level judgments can be driven by input-agnostic language priors, formatting conventions, or stereotyped reasoning templates rather than task-specific evidence. We refer to such optimization-rele
Key takeaways
- arXiv:2608.03632v1 Announce Type: new Abstract: On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals.
- Recent selective OPD methods improve this process by prioritizing signals that are confident, informative, or learnable.
- However, the assumptions overlook a fundamental failure mode of language models: their token-level judgments can be driven by input-agnostic language priors, formatting conventions, or stereotyped reasoning templates rather than task-specific evidence.
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation” may reshape data collection, model training, output accountability and market access.

Member comments