ROSS: Relearning from Self-Generated Rollouts through Selective Supervision
Quick summary
arXiv:2609.35954v1 Announce Type: cross Abstract: Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollo
Key takeaways
- arXiv:2609.35954v1 Announce Type: cross Abstract: Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances.
- Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably.
- However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision.
Why it matters
“ROSS: Relearning from Self-Generated Rollouts through Selective Supervision” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Member comments