Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack
Quick summary
arXiv:2609.08966v1 Announce Type: new Abstract: Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subsequent training. We show that this assumption can fail in a full 30B mixture-of-experts training pipeline. The checkpoints that perform better after the full downstream training stack also have higher solution density, i.e., retain downstream performance under local weight perturbations.
Key takeaways
- arXiv:2609.08966v1 Announce Type: new Abstract: Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subsequent training.
- We show that this assumption can fail in a full 30B mixture-of-experts training pipeline.
- The checkpoints that perform better after the full downstream training stack also have higher solution density, i.e., retain downstream performance under local weight perturbations.
Why it matters
“Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments