arXiv Artificial Intelligence

OSGuard: A Benchmark for Safety in Computer-Use Agents

OSGuard: A Benchmark for Safety in Computer-Use Agents

Quick summary

arXiv:2606.15034v2 Announce Type: replace Abstract: Computer-use agents can complete benign user instructions while violating important constraints of the user's environment. We introduce OSGuard, a dual-granularity benchmark suite for evaluating safety through local, pre-execution guardrail decisions and end-to-end task execution. Its action-level benchmark contains 324 human-annotated examples in which guardrails classify candidate actions as allowed, unrelated, or unsafe given the original instruction and current interface state. Its risk-augmented execution suite contains 45 tasks derived

Key takeaways

  • arXiv:2606.15034v2 Announce Type: replace Abstract: Computer-use agents can complete benign user instructions while violating important constraints of the user's environment.
  • We introduce OSGuard, a dual-granularity benchmark suite for evaluating safety through local, pre-execution guardrail decisions and end-to-end task execution.
  • Its action-level benchmark contains 324 human-annotated examples in which guardrails classify candidate actions as allowed, unrelated, or unsafe given the original instruction and current interface state.

Why it matters

This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗