arXiv Artificial Intelligence

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

Quick summary

arXiv:2605.11723v3 Announce Type: replace-cross Abstract: In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global temporal scan to anchor anomalous time windows, then performs fine-grained spatial grounding within the localized interval, and finally derives robust judgments via structured spatiotemporal Chain-of-Thought reasoning. To equip the model with these capabilities, we construct the first large-scale generated video anomaly dataset with per-frame bounding-box annotat

Key takeaways

  • arXiv:2605.11723v3 Announce Type: replace-cross Abstract: In this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models.
  • During inference, it first conducts a global temporal scan to anchor anomalous time windows, then performs fine-grained spatial grounding within the localized interval, and finally derives robust judgments via structured spatiotemporal Chain-of-Thought reasoning.
  • To equip the model with these capabilities, we construct the first large-scale generated video anomaly dataset with per-frame bounding-box annotat

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗