CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
Kechen Liu, Ola Shorinwa
arXiv:2608.27406v1Today’s pick is CLAP, a new video world model that tries to learn physics across different kinds of bodies, not just one robot. The problem it tackles is a big one in robotics: most action-conditioned video models only work for a single embodiment, so they can’t fully use the huge amount of human and robot video available on the internet. CLAP’s core idea is to align these different action spaces by using end-effector poses, language instructions, and even latent actions, then train in a curriculum that first learns broad physical priors from unlabeled video and later grounds them for real robot control. That matters because it turns diverse video into a shared simulator for prediction and planning. In experiments, CLAP matches or beats strong single-robot baselines and can transfer zero-shot to new tasks and robot types. If this scales, it could make robot learning far more data-efficient and much more general.
Also spotted that day
Previous daily papers