Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search
Quick summary
arXiv:2609.19799v1 Announce Type: cross Abstract: LLM-driven evolutionary search finds programs by launching seeds and iterating each one. Papers report a single budget setting, usually one seed run for a fixed number of iterations, and rank methods from that one point. We show this is not enough. We evaluate three evolutionary search strategies on five optimization tasks, commonly used by papers in the genre to report results. We run the analysis over a full grid of seeds and iterations. Our findings suggest that the best way to split a fixed budget between more seeds (width) and more iterati
Key takeaways
- arXiv:2609.19799v1 Announce Type: cross Abstract: LLM-driven evolutionary search finds programs by launching seeds and iterating each one.
- Papers report a single budget setting, usually one seed run for a fixed number of iterations, and rank methods from that one point.
- We evaluate three evolutionary search strategies on five optimization tasks, commonly used by papers in the genre to report results.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments