Trust the Critic More
Quick summary
arXiv:2609.39247v1 Announce Type: cross Abstract: Standard language model RL algorithms credit every token of a long rollout with the same advantage determined by the terminal reward. Actor-critic methods can provide finer-grained credit assignment, but learned critics are generally considered too inaccurate to trust when training LLMs with RL. In recent works, even when a critic is present, it is used only for baseline estimation, so every trajectory must be rolled out to its terminal reward. We introduce Actor-Critic with Action Chunking (AC2) that removes the need to roll every trajectory t
Key takeaways
- arXiv:2609.39247v1 Announce Type: cross Abstract: Standard language model RL algorithms credit every token of a long rollout with the same advantage determined by the terminal reward.
- Actor-critic methods can provide finer-grained credit assignment, but learned critics are generally considered too inaccurate to trust when training LLMs with RL.
- In recent works, even when a critic is present, it is used only for baseline estimation, so every trajectory must be rolled out to its terminal reward.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments