DemoEvolve: Demonstration-Guided Harness Evolution under Sparse Feedback
Quick summary
arXiv:2605.24539v2 Announce Type: replace Abstract: Harness evolution enables frozen language model agents to adapt to unfamiliar tasks by modifying the external programs that govern their behavior. For long-horizon tasks, each rollout is costly, and a limited interaction budget may yield few examples of effective behavior. Sparse, delayed feedback also makes it difficult for a coding agent to diagnose failures and determine which modifications will improve performance. We present DemoEvolve, which uses human demonstrations to guide harness evolution under limited interaction budgets. The codi
Key takeaways
- arXiv:2605.24539v2 Announce Type: replace Abstract: Harness evolution enables frozen language model agents to adapt to unfamiliar tasks by modifying the external programs that govern their behavior.
- For long-horizon tasks, each rollout is costly, and a limited interaction budget may yield few examples of effective behavior.
- Sparse, delayed feedback also makes it difficult for a coding agent to diagnose failures and determine which modifications will improve performance.
Why it matters
“DemoEvolve: Demonstration-Guided Harness Evolution under Sparse Feedback” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments