arXiv Artificial IntelligenceWhen Words Are Safe But Actions Kill: Probing Physical Jailbreak Beyond Textual Jailbreak in Hidden-State Risk Space
arXiv:2607.15218v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly serve as…
