arXiv Artificial Intelligence

CruxBench: A Benchmark of Information Discovery

CruxBench: A Benchmark of Information Discovery

Quick summary

arXiv:2609.35879v1 Announce Type: cross Abstract: Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whose answers provide key steps on the path toward solving the target problem. To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information

Key takeaways

  • arXiv:2609.35879v1 Announce Type: cross Abstract: Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels.
  • But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whose answers provide key steps on the path toward solving the target problem.
  • To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information

Why it matters

“CruxBench: A Benchmark of Information Discovery” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗