Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World Model
Xiwen Chen, Rigaudiere Z. Li, Zhiruo Zhou, Xiaojun Zhu, Houde Liu
arXiv:2609.28414v1Latent world models are a promising way to predict how scenes change over time, but there’s a subtle failure mode: they can become very good at predicting static appearance while losing the motion that matters for control. This paper studies that problem in frozen self-supervised latent spaces, where training is cheap and stable but the manipulated object barely moves in the model’s predictions. The authors trace the issue to the supervision signal itself: if the model only sees latent losses, it is never told where motion should happen along the rollout. Their fix is Decode-Augmented Rollout Training, or DART, which keeps the representation frozen and retrains only the flow model with decode-path supervision. That extra signal restores temporal structure, makes predicted motion line up with the scene, and improves performance while preserving the efficiency of the original setup. The work also exposes a useful evaluation warning: pixel error alone can make a frozen, motionless predictor look better than it really is.
Also spotted that day
Previous daily papers