Evaluating Large Language Models for Symbolic Security Protocol Analysis
Quick summary
arXiv:2607.20712v2 Announce Type: replace-cross Abstract: Security protocols verification relies on formal tools such as ProVerif and OFMC. This study evaluates whether large language models (LLMs) can perform comparable analysis. We test GPT and DeepSeek in chat and reasoning modes over three runs on 130 obfuscated AnB/AnBx protocols covering 388 security goals, scored against ProVerif and OFMC. Each provider uses a single model in both modes, switching reasoning on and off, so both contrasts isolate reasoning itself. Chat models achieve 72.7% recall at 27.3% precision for GPT and 69.3% recal
Key takeaways
- arXiv:2607.20712v2 Announce Type: replace-cross Abstract: Security protocols verification relies on formal tools such as ProVerif and OFMC.
- This study evaluates whether large language models (LLMs) can perform comparable analysis.
- We test GPT and DeepSeek in chat and reasoning modes over three runs on 130 obfuscated AnB/AnBx protocols covering 388 security goals, scored against ProVerif and OFMC.
Why it matters
“Evaluating Large Language Models for Symbolic Security Protocol Analysis” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Member comments