Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English
Quick summary
arXiv:2503.17039v3 Announce Type: replace-cross Abstract: Automatic text summarization relies on automatic evaluation to quickly determine the quality of summarization models via automatic metrics and LLM-as-a-Judge models. However, these techniques require meta-evaluation to ensure that they capture human judgments correctly. In this paper, we explore this meta-evaluation beyond English by generating a new multilingual summary meta-evaluation dataset (BASSE), which comprises human judgments on 2,040 abstractive summaries, generated either manually or by five Large Language Models (LLMs) with
Key takeaways
- arXiv:2503.17039v3 Announce Type: replace-cross Abstract: Automatic text summarization relies on automatic evaluation to quickly determine the quality of summarization models via automatic metrics and LLM-as-a-Judge models.
- However, these techniques require meta-evaluation to ensure that they capture human judgments correctly.
- In this paper, we explore this meta-evaluation beyond English by generating a new multilingual summary meta-evaluation dataset (BASSE), which comprises human judgments on 2,040 abstractive summaries, generated either manually or by five Large Language Models (LLMs) with
Why it matters
“Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments