arXiv Artificial Intelligence

Selective Off-Policy Reference Tuning with Plan Guidance

Selective Off-Policy Reference Tuning with Plan Guidance

Quick summary

arXiv:2605.11505v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference solution, compares token probabilities with and without that plan, and gives higher weight to tokens that become more predictable under plan conditioning. This turns all-wrong prompts into selective, structure-aware learning signals instead of uniform imitation. Across three backbones a

Key takeaways

  • arXiv:2605.11505v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail.
  • SORT adds a repair update for those failures without changing rollout generation: it derives a plan from the reference solution, compares token probabilities with and without that plan, and gives higher weight to tokens that become more predictable under plan conditioning.
  • This turns all-wrong prompts into selective, structure-aware learning signals instead of uniform imitation.

Why it matters

“Selective Off-Policy Reference Tuning with Plan Guidance” may affect what data AI products can use and where accountability sits. Product teams should watch compliance duties, rights holders should watch enforcement, and users should watch transparency and appeal mechanisms.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗