CoachWorld — Pretrained Robot World Model

Paper · Code · Project Page & Videos

CoachWorld predicts future robot observations conditioned on visual history and robot actions. It is the world-model component of RoboCoach: World Models as Active Coaches for Compositional Robot Skills.

This repository hosts the CoachWorld pretrained checkpoint, trained on heterogeneous single-arm and dual-arm robot data.

Overview

CoachWorld is an action-conditioned video world model built on an adapted Wan2.2 TI2V-5B backbone. It uses visual history and end-effector action conditioning to generate future observations, with camera and action conventions that support learning across robot datasets.

The model supports autoregressive rollout: generated observations become part of the context for subsequent predictions. In the full RoboCoach framework, skill policies interact with the world model, while a separate progress judge identifies subtask failures. These diagnoses guide targeted demonstration collection and policy updates.

Model at a Glance

Item Description
Release CoachWorld pretrain
Model type Action-conditioned video world model
Backbone Adapted Wan2.2
Conditioning Visual history and robot end-effector actions, using the prescribed camera/action representation
Output Predicted future RGB observations
Rollout mode Autoregressive video prediction
Training scope Heterogeneous single-arm and dual-arm robot data
Implementation Custom coachworld Python package

Getting Started

Download the weights from the Files and versions tab and use the CoachWorld code repository for environment setup and model integration.

The implementation includes:

  • Model conditioning, inference, and training components.
  • Modified Wan2.2 model and VAE components.
  • Video-latent data interfaces and action/camera conventions.
  • Camera and robot geometry helpers.
  • A distributed training entry point and example configuration.

Input preparation matters. Robot actions must follow the implementation’s coordinate-frame, normalization, and temporal-sampling conventions. Raw joint commands or pixel-space trajectories are not interchangeable with the model’s expected conditioning.

Use the CoachWorld implementation to load and run this checkpoint; this release does not claim compatibility with a generic Transformers or Diffusers loading pipeline.

Citation

If you use CoachWorld in your research, please cite:

@article{liu2026robocoach,
  title={RoboCoach: World Models as Active Coaches for Compositional Robot Skills},
  author={Liu, Jiajun and Chen, Yifan and Liu, Yichao and Zhang, Jiayi
          and Chen, Ruoqu and Xie, Shaoxuan and Yao, Guocai
          and Xu, Mengdi and Cui, Sen and Zhang, Changshui},
  journal={arXiv preprint arXiv:2609.39685},
  year={2026},
  url={https://arxiv.org/abs/2609.39685}
}

Acknowledgments

CoachWorld builds on Wan2.2 and other open-source components. See the code repository’s third-party notices for component attribution and applicable terms.

Contact

For implementation questions and reproducibility issues, please open an issue in the CoachWorld repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Paper for JEdward/CoachWorld