arXiv Artificial Intelligence

TSQueryBench: LLM-as-a-Judge for Time Series Explanations

TSQueryBench: LLM-as-a-Judge for Time Series Explanations

Quick summary

arXiv:2604.02118v2 Announce Type: replace Abstract: Natural language explanations of time series data are increasingly produced by foundation models in high stakes domains, making factual correctness critical. Evaluating such explanations differs fundamentally from standard natural language generation: correctness requires verifying numerical claims against structured data rather than similarity to reference text. While LLM as a Judge has emerged as a scalable paradigm for text evaluation, its applicability to numerically grounded time series explanations remains unstudied. We introduce TSQuer

Key takeaways

  • arXiv:2604.02118v2 Announce Type: replace Abstract: Natural language explanations of time series data are increasingly produced by foundation models in high stakes domains, making factual correctness critical.
  • Evaluating such explanations differs fundamentally from standard natural language generation: correctness requires verifying numerical claims against structured data rather than similarity to reference text.
  • While LLM as a Judge has emerged as a scalable paradigm for text evaluation, its applicability to numerically grounded time series explanations remains unstudied.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗