PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning
Quick summary
arXiv:2605.09860v5 Announce Type: replace Abstract: Long-horizon reasoning requires deciding not only what actions to take, but how many to execute open-loop before replanning. This number, the execution depth, balances replanning cost against compounding execution errors. Most current systems either fix the execution depth as a hand-tuned scalar or adjust it at inference time with heuristic rules decoupled from the policy; we argue both can be suboptimal. In this work, we treat the execution depth as a learnable, history-conditioned variable of the policy itself, and propose PACE, a model-nat
Key takeaways
- arXiv:2605.09860v5 Announce Type: replace Abstract: Long-horizon reasoning requires deciding not only what actions to take, but how many to execute open-loop before replanning.
- This number, the execution depth, balances replanning cost against compounding execution errors.
- Most current systems either fix the execution depth as a hand-tuned scalar or adjust it at inference time with heuristic rules decoupled from the policy; we argue both can be suboptimal.
Why it matters
The significance is not only the legal text but how it changes product design. Decisions around “PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning” may reshape data collection, model training, output accountability and market access.

Member comments