arXiv Artificial Intelligence

Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction

Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction

Quick summary

arXiv:2609.21437v1 Announce Type: cross Abstract: We propose LoG-VGGT, a memory-efficient framework for long-sequence 3D reconstruction that balances local temporal modeling with global camera consistency. Instead of relying on full global attention, our method introduces cross-window attention at a small subset of transformer blocks, enabling effective information propagation across adjacent temporal windows while keeping memory usage bounded. To mitigate long-term pose drift, we further design a global camera consistency refinement module, where camera tokens interact with compact register t

Key takeaways

  • arXiv:2609.21437v1 Announce Type: cross Abstract: We propose LoG-VGGT, a memory-efficient framework for long-sequence 3D reconstruction that balances local temporal modeling with global camera consistency.
  • Instead of relying on full global attention, our method introduces cross-window attention at a small subset of transformer blocks, enabling effective information propagation across adjacent temporal windows while keeping memory usage bounded.
  • To mitigate long-term pose drift, we further design a global camera consistency refinement module, where camera tokens interact with compact register t

Why it matters

The importance of “Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction” will be measured by what changes in practice. User behavior, access conditions, verifiable performance and responsible-use outcomes are the signals worth following.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗