arXiv Artificial Intelligence

ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

Quick summary

arXiv:2609.35954v1 Announce Type: cross Abstract: Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollo

Key takeaways

  • arXiv:2609.35954v1 Announce Type: cross Abstract: Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances.
  • Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably.
  • However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision.

Why it matters

“ROSS: Relearning from Self-Generated Rollouts through Selective Supervision” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗