Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection
Quick summary
arXiv:2609.04533v1 Announce Type: cross Abstract: Prompt injection is widely recognized as a major security threat to AI agents that interact with untrusted external data, such as websites, documents, and emails. Prior work has shown that, in the text domain, black-box prompt injection can achieve near-perfect attack success rates (ASRs). In the image domain, however, existing visual prompt injection methods are substantially less effective in attacking frontier commercial VLMs for materially harmful behavior. Achieving such outputs is hard because it requires a long and/or format-compliant ta
Key takeaways
- arXiv:2609.04533v1 Announce Type: cross Abstract: Prompt injection is widely recognized as a major security threat to AI agents that interact with untrusted external data, such as websites, documents, and emails.
- Prior work has shown that, in the text domain, black-box prompt injection can achieve near-perfect attack success rates (ASRs).
- In the image domain, however, existing visual prompt injection methods are substantially less effective in attacking frontier commercial VLMs for materially harmful behavior.
Why it matters
“Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Member comments