Defensive Sufficiency in a Stackelberg Model of AI Security
Quick summary
arXiv:2610.09892v2 Announce Type: replace-cross Abstract: Feedback from automated testing, human red teaming, and incident response can strengthen an AI system's defenses when discovered failures lead to effective repairs. We study when this feedback process provides sufficient protection and when investing in it is economically worthwhile. We begin by showing that an attack surface composed of finite number of inputs is defended with probability 1 if every unresolved attack has a persistent chance of discovery, repairs are effective, and subsequent updates preserve earlier protection. We deri
Key takeaways
- arXiv:2610.09892v2 Announce Type: replace-cross Abstract: Feedback from automated testing, human red teaming, and incident response can strengthen an AI system's defenses when discovered failures lead to effective repairs.
- We study when this feedback process provides sufficient protection and when investing in it is economically worthwhile.
- We begin by showing that an attack surface composed of finite number of inputs is defended with probability 1 if every unresolved attack has a persistent chance of discovery, repairs are effective, and subsequent updates preserve earlier protection.
Why it matters
“Defensive Sufficiency in a Stackelberg Model of AI Security” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Member comments