CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
Paper Guide Brief
Reading Brief
CLAP introduces a framework for training cross-embodiment action-conditioned video world models that learn generalizable physical priors from diverse human and robot videos, using a curriculum that combines latent actions with end-effector actions for zero-shot deployment and few-shot adaptation to new embodiments.
Central Claim
A curriculum-based cross-embodiment learning recipe that first learns foundational physical priors from unlabeled video data using latent actions and then grounds them in end-effector action spaces for zero-shot deployment, plus a few-shot adaptation paradigm...
Contribution
A curriculum-based cross-embodiment learning recipe that first learns foundational physical priors from unlabeled video data using latent actions and then grounds them in end-effector action spaces for zero-shot deployment, plus a few-shot adaptation paradigm for training single-embodiment video world models.
Why It Matters
This contribution matters because it enables video world models to leverage internet-scale heterogeneous video data across embodiments, achieving zero-shot generalization and surpassing single-embodiment baselines, which could transform ho...
Prerequisites
cross-embodiment learning, action-conditioned video generation, latent action models, end-effector actions, language actions
Atlas Placement
Robot Learning (subfield)
Read If
You care about cross-embodiment learning, action-conditioned video generation, latent action models.
Skip If
You only care about SSIM, PSNR.
Noosaga Placements
- The paper focuses on learning-based methods for robot manipulation, specifically cross-embodiment video world models and few-shot adaptation, which are core topics in robot learning.CLAP, a framework for cross-embodiment action-conditioned video generationCLAP establishes a novel paradigm for training single-embodiment video world models via sample-efficient, few-shot adaptation
- Classical Deliberative Roboticsframework80%The paper is situated within learning-based robotics, as it proposes a learning framework for video world models and uses learned representations for actions.CLAP, a framework for cross-embodiment action-conditioned video generationCLAP introduces a curriculum-based cross-embodiment learning recipe
- The work is primarily in robotics, addressing cross-embodiment learning and zero-shot generalization for manipulation tasks, which falls under the AI robotics umbrella.CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actordemonstrate CLAP’s zero-shot generalization to real-world tasks
- Deep Learningframework70%The method uses deep learning models, including video diffusion models and latent action models, which are core to deep learning.CLAP uses the video diffusion paradigm for action-conditioned video generationLatent action models (LAMs) utilize autoencoders (VAEs) to learn pseudo-actions
- The method involves video generation and prediction, which are core computer vision tasks, and uses perceptual metrics for evaluation.action-conditioned video generationWe benchmark the video models with the following perceptual metrics: SSIM, PSNR, LPIPS, FVD, and FID
- Deep Generative Modelsframework60%The paper uses deep generative models, specifically video diffusion models, for future prediction.CLAP uses the video diffusion paradigm for action-conditioned video generationwe train latent video diffusion models using continuous VAEs for video tokenization
- The approach uses deep learning models, including video diffusion models and latent action models, which are central to deep learning.CLAP uses the video diffusion paradigm for action-conditioned video generationLatent action models (LAMs) utilize autoencoders (VAEs) to learn pseudo-actions
- Learning from Demonstrationframework50%The work involves learning from demonstrations, as it trains on robot and human video data to learn physical priors, though it is not the primary focus.trained on diverse, internet-scale videos across human and robotic agentslearn foundational physical priors across unlabeled video data
- The experiments focus on robot manipulation tasks such as pick-and-place and towel folding, and the models are evaluated on manipulation datasets.we demonstrate CLAP’s zero-shot generalization to real-world manipulation taskssuccess rates in inference-time cross-policy planning in single-arm manipulation
- Data-Driven and Learning-Based Manipulationframework50%The paper addresses data-driven and learning-based manipulation, as it uses video world models for manipulation tasks and policy finetuning.demonstrate CLAP’s zero-shot generalization to real-world manipulation tasksfinetuning robot policies via video model-based reinforcement learning
Abstract
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents. CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is non-trivial because action representations vary sharply across robot platforms and are typically absent in human videos. CLAP addresses this fundamental challenge through the following core contributions. First, CLAP reconciles disparate action spaces using end-effector poses, language instructions, and latent actions. Second, to resolve their individual limitations, CLAP introduces a curriculum-based cross-embodiment learning recipe that first learns foundational physical priors across unlabeled video data using latent actions and subsequently grounds them in end-effector action spaces for zero-shot deployment to real-world tasks. Crucially, CLAP approaches or surpasses state-of-the-art single-embodiment video models in challenging environments like DROID. These performance advantages compound via few-shot adaptation to establish a novel paradigm for training single-embodiment video world models. Ultimately, CLAP delivers the most comprehensive suite of action-conditioned video world models to date - spanning diverse action-conditioning spaces (end-effector, language, and latent) and robot morphologies (including cross-embodiment, DROID, Bridge, bimanual YAM robots, and G1 humanoids). We open-source all code and models. Project Website at https://omni-clap.github.io .
Paper Context
Classified from the full extracted paper text (98,120 characters). The Paper Guide brief above is the user-facing synthesis; raw context is kept out of the page.
Full-paper context sent 98,120 of 98,120 extracted characters to classification.