arXiv Artificial Intelligence

VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

Quick summary

arXiv:2608.12246v1 Announce Type: cross Abstract: Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the full range of vulnerable software versions. Existing vulnerability datasets suffer from limited programming language coverage, restricted patch complexity, and narrow project scope. Through our dual annotation by human experts and an agentic workflow, we create a benchmark - VICBench - of 100 verified VICs for 100 CVEs ac

Key takeaways

  • arXiv:2608.12246v1 Announce Type: cross Abstract: Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases.
  • VICs are essential for determining the full range of vulnerable software versions.
  • Existing vulnerability datasets suffer from limited programming language coverage, restricted patch complexity, and narrow project scope.

Why it matters

“VICBench: A Multi-Language Benchmark for Code Vulnerability Detection” shows why AI risk cannot be reduced to answer accuracy. Access controls, logging, human approval and incident response need to be designed into the workflow from the start.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗