arXiv Artificial Intelligence

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

Quick summary

arXiv:2610.01618v1 Announce Type: new Abstract: Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. We investigate fiv

Key takeaways

  • arXiv:2610.01618v1 Announce Type: new Abstract: Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness.
  • Yet they typically treat the agent itself as fixed.
  • In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs.

Why it matters

“Agents Are Systems, Not Models: Rethinking Agentic Evaluation” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗