arXiv Artificial Intelligence

MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes

MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes

Quick summary

arXiv:2609.23980v1 Announce Type: cross Abstract: AI agents now report vulnerabilities faster than maintainers can review them. Reports often depend on security properties specific to the application, and require considerable human labor to process. To mitigate this, we introduce a framework for evaluating vulnerability reports via probes, executable checks of security properties. A reported exploit is evaluated by replaying it against the application and running the probes: a triggered probe indicates both that the exploit succeeded and which security property it violated. As a probe encodes

Key takeaways

  • arXiv:2609.23980v1 Announce Type: cross Abstract: AI agents now report vulnerabilities faster than maintainers can review them.
  • Reports often depend on security properties specific to the application, and require considerable human labor to process.
  • To mitigate this, we introduce a framework for evaluating vulnerability reports via probes, executable checks of security properties.

Why it matters

“MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗