arXiv Artificial Intelligence

Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English

Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English

Quick summary

arXiv:2503.17039v3 Announce Type: replace-cross Abstract: Automatic text summarization relies on automatic evaluation to quickly determine the quality of summarization models via automatic metrics and LLM-as-a-Judge models. However, these techniques require meta-evaluation to ensure that they capture human judgments correctly. In this paper, we explore this meta-evaluation beyond English by generating a new multilingual summary meta-evaluation dataset (BASSE), which comprises human judgments on 2,040 abstractive summaries, generated either manually or by five Large Language Models (LLMs) with

Key takeaways

  • arXiv:2503.17039v3 Announce Type: replace-cross Abstract: Automatic text summarization relies on automatic evaluation to quickly determine the quality of summarization models via automatic metrics and LLM-as-a-Judge models.
  • However, these techniques require meta-evaluation to ensure that they capture human judgments correctly.
  • In this paper, we explore this meta-evaluation beyond English by generating a new multilingual summary meta-evaluation dataset (BASSE), which comprises human judgments on 2,040 abstractive summaries, generated either manually or by five Large Language Models (LLMs) with

Why it matters

“Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗