arXiv Artificial Intelligence

OVIBench: Benchmarking Online Video Question Answering under Interruption

OVIBench: Benchmarking Online Video Question Answering under Interruption

Quick summary

arXiv:2608.22279v1 Announce Type: cross Abstract: Recent vision language models (VLMs) have achieved strong progress in video understanding. However, most existing video QA research and benchmarks still follow an offline, single-round paradigm, overlooking realistic interactions where users may interrupt the model during answer generation. To address this gap, we formulate the task of Online Video Question Answering under Interruption and introduce OVIBench, the first standardized benchmark for evaluating VLMs in this setting. OVIBench categorizes interruptions into three types: Cancellation,

Key takeaways

  • arXiv:2608.22279v1 Announce Type: cross Abstract: Recent vision language models (VLMs) have achieved strong progress in video understanding.
  • However, most existing video QA research and benchmarks still follow an offline, single-round paradigm, overlooking realistic interactions where users may interrupt the model during answer generation.
  • To address this gap, we formulate the task of Online Video Question Answering under Interruption and introduce OVIBench, the first standardized benchmark for evaluating VLMs in this setting.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗