arXiv Artificial Intelligence

AgentRx: Diagnosing AI Agent Failures from Execution Trajectories

AgentRx: Diagnosing AI Agent Failures from Execution Trajectories

Quick summary

arXiv:2602.02475v2 Announce Type: replace Abstract: AI agents often fail in ways that are difficult to localize because executions are probabilistic, long-horizon, multi-agent, and mediated by noisy tool outputs. We address this gap by manually annotating failed agent runs and release a novel benchmark of 170 trajectories across 11 diverse task settings, including structured API workflows, incident management, and open-ended web/file tasks. Each trajectory is annotated with a critical failure step and a category from a grounded-theory derived, cross-domain failure taxonomy. To mitigate the hum

Key takeaways

  • arXiv:2602.02475v2 Announce Type: replace Abstract: AI agents often fail in ways that are difficult to localize because executions are probabilistic, long-horizon, multi-agent, and mediated by noisy tool outputs.
  • We address this gap by manually annotating failed agent runs and release a novel benchmark of 170 trajectories across 11 diverse task settings, including structured API workflows, incident management, and open-ended web/file tasks.
  • Each trajectory is annotated with a critical failure step and a category from a grounded-theory derived, cross-domain failure taxonomy.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗