Failure-Guided Co-Evolution of Prompts and Training Data
Quick summary
arXiv:2609.15209v1 Announce Type: cross Abstract: Automatic prompt optimization (APO) improves language-model programs by revising prompts from task feedback, yet it typically holds its training data fixed. Repeatedly optimizing against the same instances confines feedback to weaknesses already represented in those data, leaving related failure conditions unexplored. We therefore view each failure as a dual signal: it indicates both how the prompt should be revised and what new training evidence should be synthesized. We introduce FORGE, a failure-guided framework that co-evolves prompts and t
Key takeaways
- arXiv:2609.15209v1 Announce Type: cross Abstract: Automatic prompt optimization (APO) improves language-model programs by revising prompts from task feedback, yet it typically holds its training data fixed.
- Repeatedly optimizing against the same instances confines feedback to weaknesses already represented in those data, leaving related failure conditions unexplored.
- We therefore view each failure as a dual signal: it indicates both how the prompt should be revised and what new training evidence should be synthesized.
Why it matters
“Failure-Guided Co-Evolution of Prompts and Training Data” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments