arXiv Artificial Intelligence

Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces

Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces

Quick summary

arXiv:2608.03401v1 Announce Type: cross Abstract: Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this reasoning, but a shorter trace may only show that the model stopped sooner. Here, we evaluate paired runs of the same question at matched reasoning horizons across 198 GPQA Diamond and 500 MMLU-Pro questions. We test a numeric/concision prompt that announces a token limit for Qwen3-14B and the trained effort settings of gpt-oss-20b and -120b. The Qwen prompt shortens reasoning traces by 12-17%, while accuracy

Key takeaways

  • arXiv:2608.03401v1 Announce Type: cross Abstract: Large language models often reason at length before answering, increasing cost and latency.
  • Prompts and trained settings can shorten this reasoning, but a shorter trace may only show that the model stopped sooner.
  • Here, we evaluate paired runs of the same question at matched reasoning horizons across 198 GPQA Diamond and 500 MMLU-Pro questions.

Why it matters

“Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗