iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs
Quick summary
arXiv:2609.24646v1 Announce Type: cross Abstract: On-policy self-distillation fine-tuning (SDFT) learns new skills from demonstrations while reducing forgetting, but it always distils toward the full demonstration-conditioned teacher. This fixes teacher influence at the full-teacher endpoint, providing no control over how much demonstration information should be transferred at each prediction state. We introduce Information-Proximal SDFT (iSDFT), which instead treats the teacher as a budgeted source of information. At each token, iSDFT selects the distribution closest to the current student th
Key takeaways
- arXiv:2609.24646v1 Announce Type: cross Abstract: On-policy self-distillation fine-tuning (SDFT) learns new skills from demonstrations while reducing forgetting, but it always distils toward the full demonstration-conditioned teacher.
- This fixes teacher influence at the full-teacher endpoint, providing no control over how much demonstration information should be transferred at each prediction state.
- We introduce Information-Proximal SDFT (iSDFT), which instead treats the teacher as a budgeted source of information.
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs” may reshape data collection, model training, output accountability and market access.

Member comments