WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
Quick summary
arXiv:2609.24984v1 Announce Type: cross Abstract: Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations
Key takeaways
- arXiv:2609.24984v1 Announce Type: cross Abstract: Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints.
- We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose.
- The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget.
Why it matters
This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Member comments