arXiv Artificial Intelligence

The Surface You Test Is Not the Surface That Breaks

The Surface You Test Is Not the Surface That Breaks

Quick summary

arXiv:2605.30454v2 Announce Type: replace-cross Abstract: Prompt-injection benchmarks for LLM agents typically test attacks through a single injection surface and report the resulting attack success rate as a property of the model. We ask whether those robustness conclusions remain stable when the same adversarial content enters through a different part of the agent interface. Using AgentDojo, we evaluate 13 LLMs across four task suites and place a byte-identical payload either in a tool output or in the tool description. This small change produces large differences in comparative robustness:

Key takeaways

  • arXiv:2605.30454v2 Announce Type: replace-cross Abstract: Prompt-injection benchmarks for LLM agents typically test attacks through a single injection surface and report the resulting attack success rate as a property of the model.
  • We ask whether those robustness conclusions remain stable when the same adversarial content enters through a different part of the agent interface.
  • Using AgentDojo, we evaluate 13 LLMs across four task suites and place a byte-identical payload either in a tool output or in the tool description.

Why it matters

“The Surface You Test Is Not the Surface That Breaks” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗