MiniWorld / README.md
zhaoyian01's picture
Add arXiv link and BibTeX citation
bab41c6 verified
|
Raw
History Blame Contribute Delete
9.78 kB
---
library_name: pytorch
tags:
- world-model
- video-generation
- streaming-generation
- robotics
- camera-control
- diffusion
- rectified-flow
- video-dit
- droid
- realestate10k
- arxiv:2608.01127
license: apache-2.0
---
# MiniWorld
**MiniWorld: Democratizing the Training of Video World Models from Scratch**
<div align="center" style="line-height: 1;">
<a href="https://zhao-yian.github.io/MiniWorld/" target="_blank" style="display: inline-block; vertical-align: middle; margin: 2px;">
<img alt="Project Page" src="https://img.shields.io/badge/Project-Page-1f6feb?style=for-the-badge&logo=googlechrome&logoColor=white" style="display: block;"/>
</a>
<a href="https://arxiv.org/abs/2608.01127" target="_blank" style="display: inline-block; vertical-align: middle; margin: 2px;">
<img alt="arXiv" src="https://img.shields.io/badge/arXiv-2608.01127-b31b1b?style=for-the-badge&logo=arxiv&logoColor=white" style="display: block;"/>
</a>
<a href="https://github.com/zhao-yian/MiniWorld" target="_blank" style="display: inline-block; vertical-align: middle; margin: 2px;">
<img alt="GitHub" src="https://img.shields.io/badge/GitHub-Code-181717?style=for-the-badge&logo=github&logoColor=white" style="display: block;"/>
</a>
</div>
MiniWorld is a minimal and reproducible framework for training streaming video
world models from scratch. Instead of adapting a pretrained bidirectional video
generator, it directly learns causal next-state prediction with a block-causal
Video Diffusion Transformer and Rectified Flow.
The same architecture supports two control modalities:
- **DROID:** low-level robot actions for embodied world modeling.
- **RealEstate10K:** camera poses for controllable scene prediction.
This Hugging Face repository hosts the MiniWorld model checkpoints. Code,
training scripts, and evaluation utilities live in the GitHub repository.
## Qualitative Results
Each tile is a 253-frame streaming rollout from the `1B` checkpoint, generated
from a single observed frame plus the control signal.
<div align="center" style="line-height: 1;">
<a href="https://huggingface.co/zhaoyian01/MiniWorld/resolve/main/assets/demo_droid.mp4" target="_blank" style="display: inline-block; vertical-align: top; margin: 4px;">
<img alt="DROID action-conditioned rollouts" src="assets/demo_droid.webp" width="400" style="display: block;"/>
</a>
<a href="https://huggingface.co/zhaoyian01/MiniWorld/resolve/main/assets/demo_re10k.mp4" target="_blank" style="display: inline-block; vertical-align: top; margin: 4px;">
<img alt="RealEstate10K camera-conditioned rollouts" src="assets/demo_re10k.webp" width="400" style="display: block;"/>
</a>
</div>
<p align="center"><sub><b>DROID</b> action-conditioned rollouts (left) and <b>RealEstate10K</b> camera-conditioned rollouts (right), shown at 2× speed. Click either grid for the full-resolution video, or see the <a href="https://zhao-yian.github.io/MiniWorld/">project page</a> for all 100 rollouts.</sub></p>
## Model Summary
MiniWorld uses a block-causal Video Diffusion Transformer trained with Rectified
Flow in the latent space of the Wan2.2 VAE. During inference, MiniWorld performs
streaming generation with a rolling KV cache and pipelined asynchronous
denoising, enabling long-horizon generation under bounded online computation.
Key components:
- **Block-causal Video DiT** with bidirectional attention inside each chunk and
causal attention across chunks.
- **Unified conditioning** for robot actions and camera poses through AdaLN-LoRA
modulation.
- **Chunk-oriented Probability Propagation (CoPP)** for stable non-decreasing
diffusion schedules.
- **Continued long-context training** from short clips to 253-frame sequences.
- **Structured rolling KV cache** with a persistent sink and FIFO history.
- **Pipelined asynchronous denoising** for a quality-throughput trade-off at
inference time.
The complete model can be trained in several days on a single 8-GPU server.
## Released Checkpoints
Sampling requires matching the checkpoint with the corresponding dataset and
model scale.
| Dataset | Model | Status | Checkpoint |
| --- | --- | --- | --- |
| DROID | MiniWorld-0.5B | Available | [MiniWorld_0_5b_droid.pt](https://huggingface.co/zhaoyian01/MiniWorld/resolve/main/MiniWorld_0_5b_droid.pt) |
| DROID | MiniWorld-1B | Available | [MiniWorld_1b_droid.pt](https://huggingface.co/zhaoyian01/MiniWorld/resolve/main/MiniWorld_1b_droid.pt) |
| DROID | MiniWorld-3B | Coming soon | -- |
| RealEstate10K | MiniWorld-0.5B | Available | [MiniWorld_0_5b_re10k.pt](https://huggingface.co/zhaoyian01/MiniWorld/resolve/main/MiniWorld_0_5b_re10k.pt) |
| RealEstate10K | MiniWorld-1B | Available | [MiniWorld_1b_re10k.pt](https://huggingface.co/zhaoyian01/MiniWorld/resolve/main/MiniWorld_1b_re10k.pt) |
| RealEstate10K | MiniWorld-3B | Coming soon | -- |
Download a single checkpoint with:
```bash
hf download zhaoyian01/MiniWorld \
--include "MiniWorld_1b_droid.pt" \
--local-dir checkpoints/miniworld
```
## Model Configurations
`MODEL` is the identifier expected by the training and sampling scripts in the
GitHub repository.
| Model | `MODEL` | Depth | Width | Heads | Parameters |
| --- | --- | ---: | ---: | ---: | ---: |
| MiniWorld-B | `B` | 12 | 768 | 12 | 0.12B |
| MiniWorld-L | `L` | 24 | 1024 | 16 | 0.39B |
| MiniWorld-0.5B | `0.5B` | 28 | 1152 | 16 | 0.55B |
| MiniWorld-1B | `1B` | 28 | 1536 | 12 | 1B |
| MiniWorld-3B | `3B` | 32 | 2560 | 20 | 3B |
## Intended Use
MiniWorld is intended for research on streaming video world models, including:
- action-conditioned robot world modeling,
- camera-pose-conditioned scene prediction,
- long-horizon autoregressive video generation,
- temporal memory and KV-cache mechanisms,
- train-test alignment for streaming diffusion models.
MiniWorld is a research baseline and is not intended as a general-purpose
text-to-video model.
## Requirements
Inference requires the MiniWorld codebase and the pretrained Wan2.2 VAE:
- Linux with an NVIDIA CUDA GPU
- Python 3.11
- CUDA-compatible PyTorch 2.x
- FlashAttention
- Wan2.2 VAE checkpoint from `Wan-AI/Wan2.2-TI2V-5B`
Download the VAE:
```bash
hf download Wan-AI/Wan2.2-TI2V-5B \
--include "Wan2.2_VAE.pth" \
--local-dir checkpoints/wan2.2
```
## Usage
Clone the [MiniWorld codebase](https://github.com/zhao-yian/MiniWorld), install
its requirements, then download the desired checkpoint. All commands are run
from the repository root.
The default sampler uses one observed frame as initial context, eight in-flight
chunks and a 24-chunk rolling KV cache (a 64-frame active attention window), one
persistent sink frame, 100 denoising steps with classifier-free guidance at
scale 2.0, and a 64-latent-frame rollout corresponding to 253 RGB frames.
Generated videos are saved to `${SAMPLE_DIR}/pred/`.
### DROID action-conditioned generation
```bash
DATA_ROOT=/path/to/droid_lerobot \
CKPT=/path/to/MiniWorld_1b_droid.pt \
VAE_CKPT=checkpoints/wan2.2/Wan2.2_VAE.pth \
MODEL=1B \
bash scripts/sample_droid.sh
```
### RealEstate10K camera-conditioned generation
```bash
DATA_ROOT=/path/to/re10k/videos \
POSE_DIR=/path/to/re10k/poses \
CKPT=/path/to/MiniWorld_1b_re10k.pt \
VAE_CKPT=checkpoints/wan2.2/Wan2.2_VAE.pth \
MODEL=1B \
bash scripts/sample_re10k.sh
```
### Common inference controls
```bash
GPU=0 \
TOTAL_LEN=96 \
CFG_SCALE=2.0 \
SAMPLE_NUM_VIDEOS=10 \
STREAM_INFLIGHT_CHUNKS=8 \
STREAM_MAX_CACHE_CHUNKS=24 \
STREAM_SINK_SIZE=1 \
bash scripts/sample_droid.sh
```
`TOTAL_LEN` sets the rollout length in latent frames and can exceed the trained
window, since streaming keeps the attention span bounded; `TOTAL_LEN=96` yields
381 RGB frames from a 64-frame checkpoint. MiniWorld is a streaming model and
does not assume a fixed generation horizon.
### Custom camera trajectories
A RealEstate10K checkpoint can also animate a single image along a procedural
camera trajectory, without any dataset on disk:
```bash
PYTHONPATH=. python -m miniworld.sample \
--dataset re10k \
--init_image /path/to/first_frame.png \
--custom_camera_trajectory orbit_right \
--checkpoint /path/to/MiniWorld_1b_re10k.pt \
--vae_checkpoint checkpoints/wan2.2/Wan2.2_VAE.pth \
--sample_dir samples/re10k_orbit_right \
--wm_model 1B \
--total_len 64 \
--sample_num_videos 1 \
--trajectory_magnitude 3.0
```
These checkpoints are trained on raw (unnormalized) translations, so
`--trajectory_magnitude` is worth tuning: `1.0` is almost static, `3.0` is a
good default at `--total_len 64`, and values above `5.0` degrade the second half
of the rollout. Scale it with the rollout length to keep the same apparent
speed. See the GitHub README for the full list of trajectories.
## Limitations
MiniWorld is a research model trained and evaluated at modest resolution and on
limited domains. It may exhibit long-horizon drift, geometric errors, temporal
inconsistencies, and failures under out-of-distribution actions, poses, scenes,
or camera motions. It should not be used for safety-critical simulation or as a
faithful physical simulator.
## Citation
The paper is available on arXiv: [arXiv:2608.01127](https://arxiv.org/abs/2608.01127).
If you find MiniWorld useful in your research, please cite:
```bibtex
@article{zhao2026miniworld,
title = {MiniWorld: Democratizing the Training of Video World Models from Scratch},
author = {Zhao, Yian and Zheng, Ruochong and Guo, Hongcan and Yan, Yu and Zhang, Jian and Chen, Jie},
journal = {arXiv preprint arXiv:2608.01127},
year = {2026}
}
```
## License
These checkpoints are released under the Apache 2.0 license. Please also follow
the licenses and usage terms of the underlying datasets (DROID, RealEstate10K)
and of the Wan2.2 VAE.