Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools
Quick summary
arXiv:2609.05587v1 Announce Type: new Abstract: Existing evaluations of tool-using agents primarily measure whether an agent can successfully complete diverse tasks with tools. These evaluations generally assume that tools return reliable information. However, tool returns in real-world systems can be plausible yet incorrect. We investigate how agents respond to unreliable tool returns by evaluating fourteen LLMs using three tools-web search, LLM sub-agent delegation, and code execution. For each tool, we corrupt its returns and measure whether agents adopt the corrupted content in their final
Key takeaways
- arXiv:2609.05587v1 Announce Type: new Abstract: Existing evaluations of tool-using agents primarily measure whether an agent can successfully complete diverse tasks with tools.
- These evaluations generally assume that tools return reliable information.
- However, tool returns in real-world systems can be plausible yet incorrect.
Why it matters
“Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools” is a product decision that may change how people work with AI. Its value depends on task completion, correction effort and data handling—not simply the presence of a new feature.

Member comments