Self-Play Pretraining with Zero Data
Quick summary
arXiv:2609.30063v1 Announce Type: new Abstract: Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model's behalf. A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement. This would provide an effectively unbounded source of training data, limited by compute rather than human knowledge. We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision. Our procedure casts synthetic data generat
Key takeaways
- arXiv:2609.30063v1 Announce Type: new Abstract: Advances in language modeling have been driven by scaling pretraining on ever more data.
- Yet, the training data is still largely curated on the model's behalf.
- A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement.
Why it matters
“Self-Play Pretraining with Zero Data” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments