TempCloze: Can Video-LLMs Identify the Missing Middle?
Quick summary
arXiv:2609.01515v1 Announce Type: cross Abstract: Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct sam
Key takeaways
- arXiv:2609.01515v1 Announce Type: cross Abstract: Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors.
- To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs.
- Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates.
Why it matters
“TempCloze: Can Video-LLMs Identify the Missing Middle?” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments