CoachWorld — Pretrained Robot World Model
Paper · Code · Project Page & Videos
CoachWorld predicts future robot observations conditioned on visual history and robot actions. It is the world-model component of RoboCoach: World Models as Active Coaches for Compositional Robot Skills.
This repository hosts the CoachWorld pretrained checkpoint, trained on heterogeneous single-arm and dual-arm robot data.
Overview
CoachWorld is an action-conditioned video world model built on an adapted Wan2.2 TI2V-5B backbone. It uses visual history and end-effector action conditioning to generate future observations, with camera and action conventions that support learning across robot datasets.
The model supports autoregressive rollout: generated observations become part of the context for subsequent predictions. In the full RoboCoach framework, skill policies interact with the world model, while a separate progress judge identifies subtask failures. These diagnoses guide targeted demonstration collection and policy updates.
Model at a Glance
| Item | Description |
|---|---|
| Release | CoachWorld pretrain |
| Model type | Action-conditioned video world model |
| Backbone | Adapted Wan2.2 |
| Conditioning | Visual history and robot end-effector actions, using the prescribed camera/action representation |
| Output | Predicted future RGB observations |
| Rollout mode | Autoregressive video prediction |
| Training scope | Heterogeneous single-arm and dual-arm robot data |
| Implementation | Custom coachworld Python package |
Getting Started
Download the weights from the Files and versions tab and use the CoachWorld code repository for environment setup and model integration.
The implementation includes:
- Model conditioning, inference, and training components.
- Modified Wan2.2 model and VAE components.
- Video-latent data interfaces and action/camera conventions.
- Camera and robot geometry helpers.
- A distributed training entry point and example configuration.
Input preparation matters. Robot actions must follow the implementation’s coordinate-frame, normalization, and temporal-sampling conventions. Raw joint commands or pixel-space trajectories are not interchangeable with the model’s expected conditioning.
Use the CoachWorld implementation to load and run this checkpoint; this release does not claim compatibility with a generic Transformers or Diffusers loading pipeline.
Citation
If you use CoachWorld in your research, please cite:
@article{liu2026robocoach,
title={RoboCoach: World Models as Active Coaches for Compositional Robot Skills},
author={Liu, Jiajun and Chen, Yifan and Liu, Yichao and Zhang, Jiayi
and Chen, Ruoqu and Xie, Shaoxuan and Yao, Guocai
and Xu, Mengdi and Cui, Sen and Zhang, Changshui},
journal={arXiv preprint arXiv:2609.39685},
year={2026},
url={https://arxiv.org/abs/2609.39685}
}
Acknowledgments
CoachWorld builds on Wan2.2 and other open-source components. See the code repository’s third-party notices for component attribution and applicable terms.
Contact
For implementation questions and reproducibility issues, please open an issue in the CoachWorld repository.