CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection
Quick summary
arXiv:2609.13842v1 Announce Type: cross Abstract: Recent advances in speech synthesis and voice conversion have made deepfake speech increasingly realistic, making generalization to unseen spoofing attacks a critical challenge. Pretrained speech and audio models offer a promising direction for improving robustness to such unseen attacks. Self-supervised learning (SSL) models capture fine-grained, low-level acoustic characteristics, whereas Auditory Large Language Models (ALLMs) provide higher-level contextual representations. These complementary views can provide useful cues for improving gene
Key takeaways
- arXiv:2609.13842v1 Announce Type: cross Abstract: Recent advances in speech synthesis and voice conversion have made deepfake speech increasingly realistic, making generalization to unseen spoofing attacks a critical challenge.
- Pretrained speech and audio models offer a promising direction for improving robustness to such unseen attacks.
- Self-supervised learning (SSL) models capture fine-grained, low-level acoustic characteristics, whereas Auditory Large Language Models (ALLMs) provide higher-level contextual representations.
Why it matters
“CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Member comments