SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
Quick summary
arXiv:2609.36601v1 Announce Type: new Abstract: On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive d
Key takeaways
- arXiv:2609.36601v1 Announce Type: new Abstract: On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative.
- We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision.
- Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive d
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation” may reshape data collection, model training, output accountability and market access.

Member comments