arXiv Artificial Intelligence

Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation

Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation

Quick summary

arXiv:2609.22603v1 Announce Type: cross Abstract: Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near-identical values that fail to rank models effectively. We observe this saturation across three public datasets, two proprietary datasets, and multilingual settings. Motivated by this, we introduce Semantic Scaffold, an evaluation framework that extracts a hierarchical representation of fa

Key takeaways

  • arXiv:2609.22603v1 Announce Type: cross Abstract: Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality.
  • Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near-identical values that fail to rank models effectively.
  • We observe this saturation across three public datasets, two proprietary datasets, and multilingual settings.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗