arXiv Artificial Intelligence

RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

Quick summary

arXiv:2609.17985v1 Announce Type: new Abstract: AI agents are usually evaluated by whether they complete a task. In interactive service settings, a successful agent can still frustrate users by asking repeated questions, performing redundant searches, or making avoidable revisions. We introduce RideWay, an efficiency-centered benchmark for ridehailing agents in a stateful tool-calling environment, together with Efficiency Utility, a success-gated metric that discounts successful trajectories for excess tool calls and user-facing turns relative to task-specific reference effort. Human paired pr

Key takeaways

  • arXiv:2609.17985v1 Announce Type: new Abstract: AI agents are usually evaluated by whether they complete a task.
  • In interactive service settings, a successful agent can still frustrate users by asking repeated questions, performing redundant searches, or making avoidable revisions.
  • We introduce RideWay, an efficiency-centered benchmark for ridehailing agents in a stateful tool-calling environment, together with Efficiency Utility, a success-gated metric that discounts successful trajectories for excess tool calls and user-facing turns relative to task-specific reference effort.

Why it matters

This development shows AI moving deeper into everyday software. Productivity potential should be weighed against price, data permissions, exportability and the preservation of human control.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗