Chinese Competitive Debating Dataset and Benchmark
Quick summary
arXiv:2609.21637v1 Announce Type: cross Abstract: Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubr
Key takeaways
- arXiv:2609.21637v1 Announce Type: cross Abstract: Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric.
- We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels.
- We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubr
Why it matters
“Chinese Competitive Debating Dataset and Benchmark” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments