What Makes Adversarial Examples Transfer Across Deepfake Detectors?
Quick summary
arXiv:2609.10002v1 Announce Type: cross Abstract: Deepfake detectors remain vulnerable to transfer-based black-box attacks, in which adversarial examples are generated on a source surrogate model and transferred to a target model, unknown to the attacker. Yet how source--target compatibility shapes attack success remains poorly understood. Prior studies evaluate limited detector pools and rarely disentangle architectural from training factors. We conduct a controlled evaluation of adversarial transferability across 60 detectors spanning six backbones, two pretraining regimes, and five training
Key takeaways
- arXiv:2609.10002v1 Announce Type: cross Abstract: Deepfake detectors remain vulnerable to transfer-based black-box attacks, in which adversarial examples are generated on a source surrogate model and transferred to a target model, unknown to the attacker.
- Yet how source--target compatibility shapes attack success remains poorly understood.
- Prior studies evaluate limited detector pools and rarely disentangle architectural from training factors.
Why it matters
This development is a reminder to test misuse and data-leak scenarios alongside speed and quality. Trust should come from testable controls and clear failure reporting, not protection claims alone.

Member comments