Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation
Quick summary
arXiv:2609.22603v1 Announce Type: cross Abstract: Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near-identical values that fail to rank models effectively. We observe this saturation across three public datasets, two proprietary datasets, and multilingual settings. Motivated by this, we introduce Semantic Scaffold, an evaluation framework that extracts a hierarchical representation of fa
Key takeaways
- arXiv:2609.22603v1 Announce Type: cross Abstract: Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality.
- Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near-identical values that fail to rank models effectively.
- We observe this saturation across three public datasets, two proprietary datasets, and multilingual settings.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments