arXiv Artificial Intelligence

AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos

AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos

Quick summary

arXiv:2609.24487v1 Announce Type: cross Abstract: In this work, we present a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis. Unlike prior methods which first estimate dense pixel correspondences and then recover object motion from them, our method infers a structured 3D object model, including its geometry and kinematic structure, and uses this model to optimise object track estimates over time. In our optimisation loop, a Vision-Language Model (VLM) agent iteratively refines shape or generalised pose through a render-and-compare loop, combining coars

Key takeaways

  • arXiv:2609.24487v1 Announce Type: cross Abstract: In this work, we present a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis.
  • Unlike prior methods which first estimate dense pixel correspondences and then recover object motion from them, our method infers a structured 3D object model, including its geometry and kinematic structure, and uses this model to optimise object track estimates over time.
  • In our optimisation loop, a Vision-Language Model (VLM) agent iteratively refines shape or generalised pose through a render-and-compare loop, combining coars

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗