Disentangling Self-Distillation: Measuring and Modeling Acquisition and Retention
Quick summary
arXiv:2609.39494v1 Announce Type: new Abstract: Self-distillation with privileged context adapts a language model from demonstrations by letting the model, once conditioned on a reference response, teach its context-free copy token by token. Our taxonomy reveals existing methods differ along three entangled axes: (i) the rollout source (student or teacher), (ii) the teacher coupling (frozen, or an exponential moving average of the student at some coupling rate) and (iii) the KL direction (reverse or forward), yet these axes are usually studied in fixed combinations and have led to conflicting
Key takeaways
- arXiv:2609.39494v1 Announce Type: new Abstract: Self-distillation with privileged context adapts a language model from demonstrations by letting the model, once conditioned on a reference response, teach its context-free copy token by token.
- Our taxonomy reveals existing methods differ along three entangled axes: (i) the rollout source (student or teacher), (ii) the teacher coupling (frozen, or an exponential moving average of the student at some coupling rate) and (iii) the KL direction (reverse or forward), yet these axes are usually studied in fixed combinations and have led to conflicting
Why it matters
This is more than a company headline: it shows who controls infrastructure, users and data in the AI value chain. The practical effect will appear in product integration, pricing and delivered capacity.

Member comments