arXiv Artificial Intelligence

SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning

SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning

Quick summary

arXiv:2609.36552v1 Announce Type: cross Abstract: Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggre

Key takeaways

  • arXiv:2609.36552v1 Announce Type: cross Abstract: Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks.
  • Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count.
  • Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggre

Why it matters

“SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗