Introduces CLAP, a cross-embodiment action-conditioned video generation framework trained on internet-scale human and robotic videos to learn generalizable physics.
The premise of the query is incorrect; the search context does not support the description of CLAP as a "Cross-Embodiment Video World Model" or a "Zero-Shot Physical Simulator."
According to the provided sources, CLAP stands for Contrastive Latent Action Pretraining. It is a framework designed to address data scarcity in robotic manipulation by aligning visual latent spaces from human videos with proprioceptive latent action spaces from robot trajectories.
Key distinctions from the query: Method: CLAP uses contrastive learning to map video transitions onto a quantized, physically executable codebook, rather than generating video as a physical simulator. Goal: It enables skill transfer from abundant human demonstrations to robotic execution, creating Vision-Language-Action (VLA) models (specifically CLAP-NTP and CLAP-RF), not general-purpose video world models. * Performance: CLAP improves instruction following and object generalization by disentangling manipulation skills from visual noise, but it is not described as a zero-shot physical simulator in the sense of predicting physics via video generation alone (a role more closely associated with frameworks like PhysWorld in the context).
The material presents CLAP, a cross-embodiment, action-conditioned video generation framework that treats video world models as zero-shot physical simulators. It is trained on internet-scale collections of human and robotic videos, spanning different embodiments, viewpoints, and interaction styles, to learn a shared representation of how actions alter visual scenes over time. Rather than relying on a single robot body, a narrow task distribution, or an explicitly engineered physics engine, CLAP conditions future-frame prediction on an initial state and a sequence of actions, aiming to infer physically plausible dynamics that transfer across embodiments and interaction contexts.
A key contribution is the cross-embodiment training recipe, which leverages heterogeneous human and robot video to capture generalizable physical priors. The framework also contributes an action-conditioned generative model capable of rolling out future observations without task-specific fine-tuning, and it reframes video generation as an implicit simulator for physical interaction. The central insight is that large-scale visual data can encode enough causal and physical structure—such as object permanence, contact, gravity, deformation, occlusion, and affordances—to support zero-shot prediction of how environments respond to actions, even for embodiments or scenes not explicitly represented during training.
This matters because it offers a scalable alternative to hand-built simulators and embodiment-specific dynamics models. If a video world model can generate physically consistent futures from actions, it can support downstream robotics tasks such as policy evaluation, planning, data augmentation, and counterfactual reasoning, while reducing dependence on curated simulation assets or labeled interaction data. More broadly, the work positions generative video models not merely as visual predictors, but as learned physical engines that can bridge perception, action, and embodiment.