Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning
Quick summary
arXiv:2609.37066v1 Announce Type: cross Abstract: Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed. Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or memorisation. We compare three post-training paths under a common diagnostic readout: our sufficiently trained off-policy distillation trajectories, released Qwen3 off-policy-plus-on-policy distillation endpoints, and a released DeepSeek-Math endpoint trained with Group Relative Policy Optimi
Key takeaways
- arXiv:2609.37066v1 Announce Type: cross Abstract: Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed.
- Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or memorisation.
- We compare three post-training paths under a common diagnostic readout: our sufficiently trained off-policy distillation trajectories, released Qwen3 off-policy-plus-on-policy distillation endpoints, and a released DeepSeek-Math endpoint trained with Group Relative Policy Optimi
Why it matters
“Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments