arXiv Artificial Intelligence

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

Quick summary

arXiv:2608.15869v1 Announce Type: cross Abstract: Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is particularly problematic for proactive video reasoning. We ask whether models can learn to think visually during training while reasoning directly at inference. We introduce Internalized Visual Thinking (IVT), a post-training framework

Key takeaways

  • arXiv:2608.15869v1 Announce Type: cross Abstract: Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments.
  • By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is particularly problematic for proactive video reasoning.
  • We ask whether models can learn to think visually during training while reasoning directly at inference.

Why it matters

“Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗