arXiv Artificial Intelligence

iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs

iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs

Quick summary

arXiv:2609.24646v1 Announce Type: cross Abstract: On-policy self-distillation fine-tuning (SDFT) learns new skills from demonstrations while reducing forgetting, but it always distils toward the full demonstration-conditioned teacher. This fixes teacher influence at the full-teacher endpoint, providing no control over how much demonstration information should be transferred at each prediction state. We introduce Information-Proximal SDFT (iSDFT), which instead treats the teacher as a budgeted source of information. At each token, iSDFT selects the distribution closest to the current student th

Key takeaways

  • arXiv:2609.24646v1 Announce Type: cross Abstract: On-policy self-distillation fine-tuning (SDFT) learns new skills from demonstrations while reducing forgetting, but it always distils toward the full demonstration-conditioned teacher.
  • This fixes teacher influence at the full-teacher endpoint, providing no control over how much demonstration information should be transferred at each prediction state.
  • We introduce Information-Proximal SDFT (iSDFT), which instead treats the teacher as a budgeted source of information.

Why it matters

The significance is not only the legal text but how it changes product design. Decisions around “iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs” may reshape data collection, model training, output accountability and market access.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗