WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
Wangbo Yu, Kunhao Liu, Wenbo Hu, Shenghai Yuan, Chaoran Feng, Haiyang Zhou, et al.
arXiv:2609.24984v1Video world models are a key step toward AI that can imagine and interact with dynamic environments, but they often forget what they saw earlier, especially when the camera moves or the rollout gets long. This paper tackles that consistency problem. The core idea is WorldCrafter, a world model with an implicit 3D-aware memory that can be queried by viewpoint. Instead of stuffing all past observations into a fixed token budget, it learns to compress multi-view evidence into target-view-specific memory tokens, then uses that memory to guide video generation. A pose-conditioned readout module helps the model retrieve the right information for the requested camera angle, without needing explicit depth matching. The result is streaming scene exploration from a single image or text prompt with much better long-horizon consistency and camera control. That matters because believable world models are foundational for robotics, simulation, gaming, and interactive AI systems that need to preserve state over time, not just generate pretty frames.
Also spotted that day
Previous daily papers