Robotics
manipulation

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

Mengqi Zhang1, Sahil Khose1, Simar Kareer2, Yuchen Song1, Unnat Jain1, Judy Hoffman1

1University of California, Irvine    2Georgia Institute of Technology

🌐 Project Page  |  📄 Paper  |  💻 Code  |  🧩 Datasets [LIBERO | RoboCasa | YAM]

DeVA is a decoupled video-action model for robot manipulation. DeVA uses two specialized experts: a video diffusion backbone and an action expert, bridged by multi-level feature alignment and physically-salient supervision (affordance + depth). This transfers the rich spatiotemporal priors of a video generative model into an efficient robot policy while training and converging faster.

Checkpoints

For every benchmark we release two variants: a base model (action + video only) and a guidance version that adds the affordance + depth heads.

Benchmark Base model w/ physical guidance
RoboCasa robocasa_base.pt robocasa_guidance.pt
LIBERO / LIBERO-Plus libero_guidance.pt
YAM bimanual yam_base.pt yam_guidance.pt

All checkpoints are trained from the Cosmos-Predict2 2B Video2World (480p / 16fps) backbone.

Component Params Status
T5-XXL text encoder 4.86 B frozen
VAE tokenizer 127 M frozen
Video2World DiT (backbone) 1.96 B trained
Action expert 564 M trained
Affordance head 17.6 M trained
Depth head 13.4 M trained
Joint-fusion projection 5.2 M trained
Total / trainable 7.55 B / 2.56 B (33.9%)

Usage

Install the code from Mq-Zhang1/deva, then download a checkpoint into checkpoints/deva/:

hf download mengqz9/DeVA libero_guidance.pt --local-dir checkpoints/deva

or use the helper script in the repo:

python scripts/download_deva_checkpoints.py --list                          # list available
python scripts/download_deva_checkpoints.py --checkpoint libero_guidance    # one (repeatable)
python scripts/download_deva_checkpoints.py --all                           # all checkpoints

Closed-loop evaluation uses a client/server split: a server hosts the DeVA model, a client drives the simulator. See the repository README for evaluation, real-robot deployment, and training instructions.

Citation

@article{zhang2026deva,
  title   = {{DeVA}: Decoupled Video-Action Model with physical guidance for robot policy learning},
  author  = {Zhang, Mengqi and Khose, Sahil and Kareer, Simar and
             Song, Yuchen and Jain, Unnat and Hoffman, Judy},
  journal = {arXiv preprint arXiv:2607.24159},
  year    = {2026}
}

Acknowledgements

Thanks for all the excellent open-source projects: Cosmos-Predict2, LIBERO, RoboCasa, Video Depth Anything and UAD. We thank the authors of these projects.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for mengqz9/DeVA

Finetuned
(7)
this model

Datasets used to train mengqz9/DeVA

Paper for mengqz9/DeVA