|
Download README.md from kdwon/CompACT: direct link, hf CLI and curl.
- Browser
- Download file 4.4 kB
-
https://huggingface.co/kdwon/CompACT/resolve/main/README.md
- Command line
-
hf download hf://kdwon/CompACT/README.md
-
curl -L -o README.md https://huggingface.co/kdwon/CompACT/resolve/main/README.md
4.4 kB
| language: | |
| - en | |
| tags: | |
| - compact | |
| - image-tokenization | |
| - world-model | |
| - robotics | |
| - pytorch | |
| - image-to-image | |
| # CompACT — 16-token checkpoints | |
| Final 16-token checkpoints for **Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model** (CVPR 2026). | |
| [Code and setup instructions](https://github.com/kdwonn/CompACT) · [Paper](https://arxiv.org/abs/2603.05438) · [Project](https://kdwonn.github.io/CompACT) | |
| ## Included models | |
| | Directory | Model | Resolution | Checkpoint | | |
| |---|---|---|---| | |
| | `tokenizer-16-224` | CompACT, 16 tokens | 224 × 224 | `checkpoints/epoch=24-step=500000.ckpt` | | |
| | `tokenizer-16-256` | CompACT, 16 tokens | 256 × 256 | `checkpoints/epoch=24-step=500000.ckpt` | | |
| | `cdit-b-16` | CDiT-B world model | 224 × 224 | `checkpoints/latest.pth.tar` | | |
| | `cdit-l-16` | CDiT-L world model | 224 × 224 | `checkpoints/latest.pth.tar` | | |
| Both world models use **`tokenizer-16-224`**. The 256-resolution tokenizer is provided separately. These are the original full training checkpoint files, including training state; world-model inference uses the `ema` weights. Only the final checkpoint for each variant is included. Exact training steps, original experiment names, file sizes, and SHA-256 checksums are in [manifest.json](manifest.json). | |
| ## Download | |
| Install the environment following the [code repository](https://github.com/kdwonn/CompACT). Run from its root: | |
| ```python | |
| from huggingface_hub import snapshot_download | |
| snapshot_download(repo_id="kdwon/CompACT", local_dir="checkpoints/CompACT") | |
| ``` | |
| To download just one variant, include its config: | |
| ```python | |
| snapshot_download( | |
| repo_id="kdwon/CompACT", | |
| local_dir="checkpoints/CompACT", | |
| allow_patterns=["tokenizer-16-224/**", "manifest.json", "SHA256SUMS"], | |
| ) | |
| ``` | |
| ## Configuration | |
| Each variant includes `.hydra/config.yaml`, in the format expected by the code repository. Set these environment variables to your local paths: | |
| ```bash | |
| export COMPACT_CKPT_ROOT="$(pwd)/checkpoints/CompACT" | |
| export BASE_TOKENIZER_CKPT=/absolute/path/to/base-checkpoints | |
| export DATASET_PREFIX=/absolute/path/to/datasets | |
| ``` | |
| The current constructors also require these initialization files: | |
| - `$BASE_TOKENIZER_CKPT/mage_vqgan/vqgan_jax_strongaug.ckpt` | |
| - `$BASE_TOKENIZER_CKPT/dinov3/dinov3_vitb16_pretrain_lvd1689m-73cec8be.pth` | |
| See the code repository for obtaining the MAGE and DINOv3 files. `DATASET_PREFIX` must be defined even for standalone tokenizer loading because the world-model loader resolves the complete saved tokenizer configuration. | |
| World-model configs use `${oc.env:COMPACT_CKPT_ROOT}/tokenizer-16-224` and `${oc.env:DATASET_PREFIX}/nwm/{recon,sacson,scand}` instead of the original machine's paths. Adjust dataset overrides to match your local layout. | |
| ## Load a tokenizer | |
| ```bash | |
| uv run load_tokenizer_checkpoint.py "$COMPACT_CKPT_ROOT/tokenizer-16-224" --no-test | |
| ``` | |
| Use `tokenizer-16-256` for the 256-resolution variant. For image encoding, use the saved DINO normalization (`dinov2_mean` and `dinov2_std`); the tokenizer's output normalization buffers describe decoded images. | |
| ## Planning evaluation | |
| With the required navigation datasets installed: | |
| ```bash | |
| uv run bash scripts/plan.sh --nproc=1 -- \ | |
| ++exp_dir="$COMPACT_CKPT_ROOT/cdit-b-16" \ | |
| ++tokenizer_path="$COMPACT_CKPT_ROOT/tokenizer-16-224" \ | |
| ++ckp=latest | |
| ``` | |
| Replace `cdit-b-16` with `cdit-l-16` to use the larger world model. Dataset setup and evaluation options are documented in the code repository. | |
| ## Validation | |
| These files passed strict weight-loading checks in the public CompACT codebase. Both tokenizers encoded and reconstructed a real image. Both the regular and EMA weights of the world models passed forward checks, and EMA models completed image-to-predicted-image inference. Validation used PyTorch 2.6.0+cu124 and bfloat16 inference on an RTX 6000 Ada GPU. These smoke checks do not constitute a rerun of the paper's evaluation benchmarks. | |
| Verify downloaded checkpoint bytes from the download directory: | |
| ```bash | |
| sha256sum -c SHA256SUMS | |
| ``` | |
| ## Citation | |
| ```bibtex | |
| @inproceedings{kim2026planning, | |
| title={Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model}, | |
| author={Kim, Dongwon and Seo, Gawon and Lee, Jinsung and Cho, Minsu and Kwak, Suha}, | |
| booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, | |
| year={2026} | |
| } | |
| ``` | |