Co-Evolving Agents: Learning from Failures as Hard Negatives
Quick summary
arXiv:2511.22254v5 Announce Type: replace Abstract: Self-evolving agents improve their performance on long-horizon tasks by learning from their own interactions with an environment. A common approach uses the resulting failed trajectories as negatives for preference training. However, collecting an agent's own failures does not ensure that they provide informative supervision for further improvement. Obvious failures may be easy to reject without learning to identify errors in more plausible attempts. These plausible but incorrect trajectories can serve as hard negatives, helping agents learn
Key takeaways
- arXiv:2511.22254v5 Announce Type: replace Abstract: Self-evolving agents improve their performance on long-horizon tasks by learning from their own interactions with an environment.
- A common approach uses the resulting failed trajectories as negatives for preference training.
- However, collecting an agent's own failures does not ensure that they provide informative supervision for further improvement.
Why it matters
“Co-Evolving Agents: Learning from Failures as Hard Negatives” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Member comments