arXiv Artificial Intelligence

CS-Guard: Benchmarking LLM Guardrails for Code Generation Security

CS-Guard: Benchmarking LLM Guardrails for Code Generation Security

Quick summary

arXiv:2609.09798v1 Announce Type: cross Abstract: Large language models (LLMs) have been ex- ploited to generate malware, but the effective- ness of guardrails for code generation secu- rity remains unclear. We introduce CS-Guard, the first benchmark to systematically evalu- ate guardrails for code generation security. It covers 1) text-to-code generation with 1000 high-quality malware-generation prompts, 7 jailbreak attacks, and a novel fictional scenario attack (FSA) that embeds malicious intent in a legitimate fictional software-development sce- nario; and 2) code-to-code generation with 33

Key takeaways

  • arXiv:2609.09798v1 Announce Type: cross Abstract: Large language models (LLMs) have been ex- ploited to generate malware, but the effective- ness of guardrails for code generation secu- rity remains unclear.
  • We introduce CS-Guard, the first benchmark to systematically evalu- ate guardrails for code generation security.
  • It covers 1) text-to-code generation with 1000 high-quality malware-generation prompts, 7 jailbreak attacks, and a novel fictional scenario attack (FSA) that embeds malicious intent in a legitimate fictional software-development sce- nario; and 2) code-to-code generation with 33

Why it matters

“CS-Guard: Benchmarking LLM Guardrails for Code Generation Security” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗