CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
Quick summary
arXiv:2608.24794v1 Announce Type: new Abstract: Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound. Treating corrective feedback as a learned in-trajectory intervention couples the two roles: the agent must decide when to request and use feedback, while the critic must infer useful corrections from outcome-confounded rollouts whose failure patterns shift as the agent improves. We introduce CAFE (Coupled Agent--Feedback Evolution), a framework in which
Key takeaways
- arXiv:2608.24794v1 Announce Type: new Abstract: Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound.
- Treating corrective feedback as a learned in-trajectory intervention couples the two roles: the agent must decide when to request and use feedback, while the critic must infer useful corrections from outcome-confounded rollouts whose failure patterns shift as the agent improves.
- We introduce CAFE (Coupled Agent--Feedback Evolution), a framework in which
Why it matters
The importance of “CAFE: Self-Improving Search Agents Need Co-Evolving Feedback” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Member comments