Has LLM Screening Performance Stalled in Software Engineering Systematic Reviews?
Quick summary
arXiv:2610.10633v1 Announce Type: cross Abstract: Screening in systematic reviews (SRs) is manual and time-consuming. Prior work has explored large language models (LLMs) for automating this step, but LLMs are evolving rapidly, so earlier performance claims may no longer accurately reflect their screening performance. We used an existing software engineering SR screening benchmark (SESR-Eval) as our data. We also power-sampled a new, smaller dataset (SESR-Eval-Mini) that allows evaluation at lower costs. Using this data, we evaluated eight new LLMs for screening performance. Additionally, we t
Key takeaways
- arXiv:2610.10633v1 Announce Type: cross Abstract: Screening in systematic reviews (SRs) is manual and time-consuming.
- Prior work has explored large language models (LLMs) for automating this step, but LLMs are evolving rapidly, so earlier performance claims may no longer accurately reflect their screening performance.
- We used an existing software engineering SR screening benchmark (SESR-Eval) as our data.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments