arXiv Artificial Intelligence

HazardAuditor: From Executable Threats to Safer Computer-Use Agents

HazardAuditor: From Executable Threats to Safer Computer-Use Agents

Quick summary

arXiv:2609.15134v1 Announce Type: new Abstract: Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing safety risks that emerge through runtime behavior rather than generated content alone. Existing guard models target static prompts and responses and are poorly suited to agent execution; existing executable safety platforms produce evaluation verdicts rather than the normalized supervision a guard model needs to learn across heterogeneous agent frameworks. We introduce HazardAuditor, an execution-grounded framework that closes both

Key takeaways

  • arXiv:2609.15134v1 Announce Type: new Abstract: Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing safety risks that emerge through runtime behavior rather than generated content alone.
  • Existing guard models target static prompts and responses and are poorly suited to agent execution; existing executable safety platforms produce evaluation verdicts rather than the normalized supervision a guard model needs to learn across heterogeneous agent frameworks.
  • We introduce HazardAuditor, an execution-grounded framework that closes both

Why it matters

“HazardAuditor: From Executable Threats to Safer Computer-Use Agents” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗