arXiv Artificial Intelligence

Chinese Competitive Debating Dataset and Benchmark

Chinese Competitive Debating Dataset and Benchmark

Quick summary

arXiv:2609.21637v1 Announce Type: cross Abstract: Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubr

Key takeaways

  • arXiv:2609.21637v1 Announce Type: cross Abstract: Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric.
  • We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels.
  • We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubr

Why it matters

“Chinese Competitive Debating Dataset and Benchmark” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗