Unifying Video Tasks via Spatiotemporal Analogy
Quick summary
arXiv:2609.33935v2 Announce Type: replace-cross Abstract: Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion. Evaluated on a diverse task taxonomy with a strict train-test sp
Key takeaways
- arXiv:2609.33935v2 Announce Type: replace-cross Abstract: Adapting video models to new tasks typically requires dedicated data curation and fine-tuning.
- While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain.
- To explore whether analogy-based methods can unify diverse video tasks and generalize to out-of-distribution scenarios, we introduce ViGeo, a framework that extends visual in-context learning to the video domain via spatiotemporal canvas completion.
Why it matters
“Unifying Video Tasks via Spatiotemporal Analogy” illustrates how changes in the AI ecosystem can affect products, workflows and user expectations together. Its lasting significance depends on measurable adoption, cost and safety outcomes.

Member comments