arXiv Artificial Intelligence

FastKernels: Benchmarking GPU Kernel Generation in Production

FastKernels: Benchmarking GPU Kernel Generation in Production

Quick summary

arXiv:2605.23215v2 Announce Type: replace-cross Abstract: LLM-based agents for GPU kernel generation are advancing rapidly, but the benchmarks they optimize against evaluate kernels in isolation, with synthetic inputs and weak baselines, rewarding sandbox speedups that break or vanish in real inference systems. We introduce FastKernels, a benchmark of 384 tasks drawn from 47 representative architectures across 8 categories, whose kernels suffice to reimplement 94.6% (472/499) of HuggingFace Transformers architectures with outputs matching the native implementations. Each task mirrors the inter

Key takeaways

  • arXiv:2605.23215v2 Announce Type: replace-cross Abstract: LLM-based agents for GPU kernel generation are advancing rapidly, but the benchmarks they optimize against evaluate kernels in isolation, with synthetic inputs and weak baselines, rewarding sandbox speedups that break or vanish in real inference systems.
  • We introduce FastKernels, a benchmark of 384 tasks drawn from 47 representative architectures across 8 categories, whose kernels suffice to reimplement 94.6% (472/499) of HuggingFace Transformers architectures with outputs matching the native implementations.

Why it matters

AI progress is not only a software story. Chips, data centers and energy decisions help determine which models can operate economically and what end users ultimately pay.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗