arXiv Artificial Intelligence

Has LLM Screening Performance Stalled in Software Engineering Systematic Reviews?

Has LLM Screening Performance Stalled in Software Engineering Systematic Reviews?

Quick summary

arXiv:2610.10633v1 Announce Type: cross Abstract: Screening in systematic reviews (SRs) is manual and time-consuming. Prior work has explored large language models (LLMs) for automating this step, but LLMs are evolving rapidly, so earlier performance claims may no longer accurately reflect their screening performance. We used an existing software engineering SR screening benchmark (SESR-Eval) as our data. We also power-sampled a new, smaller dataset (SESR-Eval-Mini) that allows evaluation at lower costs. Using this data, we evaluated eight new LLMs for screening performance. Additionally, we t

Key takeaways

  • arXiv:2610.10633v1 Announce Type: cross Abstract: Screening in systematic reviews (SRs) is manual and time-consuming.
  • Prior work has explored large language models (LLMs) for automating this step, but LLMs are evolving rapidly, so earlier performance claims may no longer accurately reflect their screening performance.
  • We used an existing software engineering SR screening benchmark (SESR-Eval) as our data.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗