CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG
Quick summary
arXiv:2605.11611v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG) systems from outcome-only supervision. Most existing methods optimize policies from uniformly sampled rollouts, implicitly treating all trajectories as equally informative. However, trajectories differ substantially in search depth and are therefore not equally informative: deeper-search trajectories contain more retrieval decision points and provide denser direct supervision for the retrieval sub
Key takeaways
- arXiv:2605.11611v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG) systems from outcome-only supervision.
- Most existing methods optimize policies from uniformly sampled rollouts, implicitly treating all trajectories as equally informative.
- However, trajectories differ substantially in search depth and are therefore not equally informative: deeper-search trajectories contain more retrieval decision points and provide denser direct supervision for the retrieval sub
Why it matters
The importance of “CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments