arXiv Artificial Intelligence

"LLM Agent Performance" Is Not a Single Evaluation Target

"LLM Agent Performance" Is Not a Single Evaluation Target

Quick summary

arXiv:2602.03238v3 Announce Type: replace Abstract: LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget. Unified execution controls these non-model factors by evaluating candidate models under the same configuration, making observed differences more attributable to the models themselves. However, model comparison is only one use of agent benchmarks. Other evaluations compare complete agent systems or test whether a fixed model or system remains stable across predeclared changes in its operating conditions. Thes

Key takeaways

  • arXiv:2602.03238v3 Announce Type: replace Abstract: LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget.
  • Unified execution controls these non-model factors by evaluating candidate models under the same configuration, making observed differences more attributable to the models themselves.
  • However, model comparison is only one use of agent benchmarks.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗