Memorization Diagnostics for Code LLMs Should be Scale-Aware
Quick summary
arXiv:2608.12771v1 Announce Type: cross Abstract: The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current literature frequently reports widespread memorization, evaluating the underlying probing techniques across dense architectures reveals a severe breakdown in their utility at scale. Traditional encoder-style probes using perturbations such as synonym fuzzing or dead-code insertion struggle to expose memorization in scaled models, even on known-contaminated benchmarks, and decoder-style probes that rely on log p
Key takeaways
- arXiv:2608.12771v1 Announce Type: cross Abstract: The extent to which large language models for code rely on memorization over genuine understanding remains highly debated.
- While current literature frequently reports widespread memorization, evaluating the underlying probing techniques across dense architectures reveals a severe breakdown in their utility at scale.
- Traditional encoder-style probes using perturbations such as synonym fuzzing or dead-code insertion struggle to expose memorization in scaled models, even on known-contaminated benchmarks, and decoder-style probes that rely on log p
Why it matters
“Memorization Diagnostics for Code LLMs Should be Scale-Aware” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Member comments