arXiv Artificial Intelligence

TempCloze: Can Video-LLMs Identify the Missing Middle?

TempCloze: Can Video-LLMs Identify the Missing Middle?

Quick summary

arXiv:2609.01515v1 Announce Type: cross Abstract: Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct sam

Key takeaways

  • arXiv:2609.01515v1 Announce Type: cross Abstract: Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors.
  • To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs.
  • Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates.

Why it matters

“TempCloze: Can Video-LLMs Identify the Missing Middle?” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗