arXiv Artificial Intelligence

K-Bench: measuring model performance on real scientific agent requests

K-Bench: measuring model performance on real scientific agent requests

Quick summary

arXiv:2608.21601v2 Announce Type: replace Abstract: Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure. Real scientific requests arrive differently. They are underspecified, they carry attachments, and they lack ground truth. We report K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs. Thre

Key takeaways

  • arXiv:2608.21601v2 Announce Type: replace Abstract: Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure.
  • Real scientific requests arrive differently.
  • They are underspecified, they carry attachments, and they lack ground truth.

Why it matters

“K-Bench: measuring model performance on real scientific agent requests” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗