A Fragility Spectrum for Recursive Language-Model Training
Quick summary
arXiv:2609.11149v1 Announce Type: cross Abstract: Model-generated text is finding its way back into training corpora, and there is plenty of evidence that training on such data over and over collapses output diversity. Prior work has studied the phenomenon itself: which protocols and which data mixtures cause collapse. But different models behave very differently under the same process. We fix one recursive contamination protocol and let 13 publicly released checkpoints form an ecosystem that shares a common corpus for five generations. The unique 4-gram outcome after five generations ranges f
Key takeaways
- arXiv:2609.11149v1 Announce Type: cross Abstract: Model-generated text is finding its way back into training corpora, and there is plenty of evidence that training on such data over and over collapses output diversity.
- Prior work has studied the phenomenon itself: which protocols and which data mixtures cause collapse.
- But different models behave very differently under the same process.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments