arXiv Artificial Intelligence

CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG

CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG

Quick summary

arXiv:2605.11611v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG) systems from outcome-only supervision. Most existing methods optimize policies from uniformly sampled rollouts, implicitly treating all trajectories as equally informative. However, trajectories differ substantially in search depth and are therefore not equally informative: deeper-search trajectories contain more retrieval decision points and provide denser direct supervision for the retrieval sub

Key takeaways

  • arXiv:2605.11611v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG) systems from outcome-only supervision.
  • Most existing methods optimize policies from uniformly sampled rollouts, implicitly treating all trajectories as equally informative.
  • However, trajectories differ substantially in search depth and are therefore not equally informative: deeper-search trajectories contain more retrieval decision points and provide denser direct supervision for the retrieval sub

Why it matters

The importance of “CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗