Research Radarcs.ROAug 27, 2026classified

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

Kechen Liu, Ola ShorinwaarXivPDF
cs.ROcs.AIcs.CV

Paper Guide Brief

Reading Brief

CLAP introduces a framework for training cross-embodiment action-conditioned video world models that learn generalizable physical priors from diverse human and robot videos, using a curriculum that combines latent actions with end-effector actions for zero-shot deployment and few-shot adaptation to new embodiments.

Central Claim

A curriculum-based cross-embodiment learning recipe that first learns foundational physical priors from unlabeled video data using latent actions and then grounds them in end-effector action spaces for zero-shot deployment, plus a few-shot adaptation paradigm...

Contribution

A curriculum-based cross-embodiment learning recipe that first learns foundational physical priors from unlabeled video data using latent actions and then grounds them in end-effector action spaces for zero-shot deployment, plus a few-shot adaptation paradigm for training single-embodiment video world models.

Why It Matters

This contribution matters because it enables video world models to leverage internet-scale heterogeneous video data across embodiments, achieving zero-shot generalization and surpassing single-embodiment baselines, which could transform ho...

Prerequisites

cross-embodiment learning, action-conditioned video generation, latent action models, end-effector actions, language actions

Atlas Placement

Robot Learning (subfield)

Read If

You care about cross-embodiment learning, action-conditioned video generation, latent action models.

Skip If

You only care about SSIM, PSNR.

Methods
cross-embodiment learningaction-conditioned video generationlatent action modelsend-effector actionslanguage actionscurriculum learningvideo diffusionfew-shot adaptation
Tasks
video predictionrobot manipulationpick-and-placetowel foldingbimanual manipulationhumanoid control
Datasets
Open X-EmbodimentEgoDexDROIDBridgeOXE-MixDreamDojo-Human
Benchmarks
SSIMPSNRLPIPSFVDFID

Noosaga Placements

  • Robot Learningsubfield90%
    The paper focuses on learning-based methods for robot manipulation, specifically cross-embodiment video world models and few-shot adaptation, which are core topics in robot learning.
    CLAP, a framework for cross-embodiment action-conditioned video generationCLAP establishes a novel paradigm for training single-embodiment video world models via sample-efficient, few-shot adaptation
  • Classical Deliberative Roboticsframework80%
    The paper is situated within learning-based robotics, as it proposes a learning framework for video world models and uses learned representations for actions.
    CLAP, a framework for cross-embodiment action-conditioned video generationCLAP introduces a curriculum-based cross-embodiment learning recipe
  • Roboticssubfield80%
    The work is primarily in robotics, addressing cross-embodiment learning and zero-shot generalization for manipulation tasks, which falls under the AI robotics umbrella.
    CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actordemonstrate CLAP’s zero-shot generalization to real-world tasks
  • Deep Learningframework70%
    The method uses deep learning models, including video diffusion models and latent action models, which are core to deep learning.
    CLAP uses the video diffusion paradigm for action-conditioned video generationLatent action models (LAMs) utilize autoencoders (VAEs) to learn pseudo-actions
  • Computer Visionsubfield70%
    The method involves video generation and prediction, which are core computer vision tasks, and uses perceptual metrics for evaluation.
    action-conditioned video generationWe benchmark the video models with the following perceptual metrics: SSIM, PSNR, LPIPS, FVD, and FID
  • Deep Generative Modelsframework60%
    The paper uses deep generative models, specifically video diffusion models, for future prediction.
    CLAP uses the video diffusion paradigm for action-conditioned video generationwe train latent video diffusion models using continuous VAEs for video tokenization
  • Deep Learningsubfield60%
    The approach uses deep learning models, including video diffusion models and latent action models, which are central to deep learning.
    CLAP uses the video diffusion paradigm for action-conditioned video generationLatent action models (LAMs) utilize autoencoders (VAEs) to learn pseudo-actions
  • Learning from Demonstrationframework50%
    The work involves learning from demonstrations, as it trains on robot and human video data to learn physical priors, though it is not the primary focus.
    trained on diverse, internet-scale videos across human and robotic agentslearn foundational physical priors across unlabeled video data
  • The experiments focus on robot manipulation tasks such as pick-and-place and towel folding, and the models are evaluated on manipulation datasets.
    we demonstrate CLAP’s zero-shot generalization to real-world manipulation taskssuccess rates in inference-time cross-policy planning in single-arm manipulation
  • Data-Driven and Learning-Based Manipulationframework50%
    The paper addresses data-driven and learning-based manipulation, as it uses video world models for manipulation tasks and policy finetuning.
    demonstrate CLAP’s zero-shot generalization to real-world manipulation tasksfinetuning robot policies via video model-based reinforcement learning

Abstract

State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents. CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is non-trivial because action representations vary sharply across robot platforms and are typically absent in human videos. CLAP addresses this fundamental challenge through the following core contributions. First, CLAP reconciles disparate action spaces using end-effector poses, language instructions, and latent actions. Second, to resolve their individual limitations, CLAP introduces a curriculum-based cross-embodiment learning recipe that first learns foundational physical priors across unlabeled video data using latent actions and subsequently grounds them in end-effector action spaces for zero-shot deployment to real-world tasks. Crucially, CLAP approaches or surpasses state-of-the-art single-embodiment video models in challenging environments like DROID. These performance advantages compound via few-shot adaptation to establish a novel paradigm for training single-embodiment video world models. Ultimately, CLAP delivers the most comprehensive suite of action-conditioned video world models to date - spanning diverse action-conditioning spaces (end-effector, language, and latent) and robot morphologies (including cross-embodiment, DROID, Bridge, bimanual YAM robots, and G1 humanoids). We open-source all code and models. Project Website at https://omni-clap.github.io .

Paper Context

Source ContextWhole paper
Budget100,000 tokens
Coverage98,120 chars

Classified from the full extracted paper text (98,120 characters). The Paper Guide brief above is the user-facing synthesis; raw context is kept out of the page.

Full-paper context sent 98,120 of 98,120 extracted characters to classification.