arXiv Artificial Intelligence

Complexity-Aware Evaluation of LLM Comprehension

Complexity-Aware Evaluation of LLM Comprehension

Quick summary

arXiv:2609.37405v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review. However, aggregate benchmark accuracy can conceal how model reliability changes as source code becomes structurally more complex. This paper presents a complexity-aware framework for evaluating LLM code comprehension using cyclomatic complexity, nesting depth, branching factor, and Halstead volume. We evaluate DeepSeek-Coder-V2 and Llam

Key takeaways

  • arXiv:2609.37405v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review.
  • However, aggregate benchmark accuracy can conceal how model reliability changes as source code becomes structurally more complex.
  • This paper presents a complexity-aware framework for evaluating LLM code comprehension using cyclomatic complexity, nesting depth, branching factor, and Halstead volume.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗