File size: 5,854 Bytes
cf20d13 0986e10 cf20d13 0986e10 c386b53 0986e10 aac0f38 0986e10 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 | ---
license: cc-by-sa-4.0
language:
- en
base_model:
- tencent/HunyuanCustom
---
<div align="center">
<h1 align="center">
CustomX: Unified Character, Action, and Scene Customization in Video World Models
</h1>
<h3 align="center">
ECCV 2026
</h3>
<p align="center">
<a href="https://arxiv.org/abs/2512.17796"><img src="https://img.shields.io/badge/arXiv-Paper-B31B1B?style=flat&labelColor=555555&logo=arxiv&logoColor=white" alt="arXiv"></a>
<a href="https://snowflakewang.github.io/CustomX_Page/"><img src="https://img.shields.io/badge/Project-Page-4F46E5?style=flat&labelColor=555555&logo=googlechrome&logoColor=white" alt="Project Page"></a>
<a href="https://github.com/snowflakewang/CustomX"><img src="https://img.shields.io/badge/Code-GitHub-181717?style=flat&labelColor=555555&logo=github&logoColor=white" alt="Code"></a>
<a href="https://huggingface.co/SnowflakeWang/CustomX"><img src="https://img.shields.io/badge/Model-Weights-FFD21E?style=flat&labelColor=555555&logo=huggingface&logoColor=FFD21E" alt="Model"></a>
</p>
</div>
---
## ๐ Abstract
Recent advances in world models have greatly enhanced interactive environment simulation. Existing methods mainly fall into two categories: (1) static world generation models, which construct 3D environments without active agents, and (2) controllable-entity models, which allow a single entity to perform limited actions in an otherwise uncontrollable environment. In this work, we introduce CustomX, leveraging the realism and structural grounding of static world generation while extending controllable-entity models to support user-specified characters capable of performing open-ended actions. Users can provide a 3DGS scene and a character, then use natural language to direct the character to perform diverse behaviors, ranging from basic locomotion to object-centric interactions, while freely exploring the environment. CustomX synthesizes temporally coherent video clips that preserve visual fidelity with the provided scene and character, formulated as a conditional autoregressive video generation problem. Built upon a pre-trained video generator, our training strategy significantly enhances motion dynamics while maintaining generalization across actions and characters. Our evaluation covers a broad range of aspects, including visual quality, character consistency, action controllability, and long-horizon coherence.
<div align="center">
<img src="./assets/teaser_v3.jpg" width="100%" alt="CustomX Teaser"/>
</div>
## ๐ ๏ธ Installation
```bash
conda create -n customx python==3.11.9
conda activate customx
pip install uv
uv pip install torch==2.4.0 torchvision==0.19.0 torchaudio==2.4.0 --index-url https://download.pytorch.org/whl/cu124
uv pip install -r requirements.txt
wget https://github.com/Dao-AILab/flash-attention/releases/download/v2.6.3/flash_attn-2.6.3+cu123torch2.4cxx11abiFALSE-cp311-cp311-linux_x86_64.whl
uv pip install flash_attn-2.6.3+cu123torch2.4cxx11abiFALSE-cp311-cp311-linux_x86_64.whl
```
## ๐ฆ Download Checkpoints
### HunyuanCustom Base Model
Download HunyuanCustom base model [here](https://huggingface.co/tencent/HunyuanCustom).
Only need to download the following files from HunyuanCustom:
```shell
CustomX
โโโ models
โโโ base
โโโ hunyuancustom_editing_720P
โ โโโ mp_rank_00_model_states.pt
โโโ vae_3d
โโโ openai_clip-vit-large-patch14
โโโ llava-llama-3-8b-v1_1
```
### CustomX LoRA
```bash
TOKEN=xxx # add your Hugging Face token here
hf download --token $TOKEN \
SnowflakeWang/CustomX \
--local-dir models/lora
```
## ๐ Inference
### Input Preparation
We provide a sample input in the [input](https://github.com/snowflakewang/CustomX/tree/main/input) directory. To use your own inputs, organize them following the same structure.
```shell
input
โโโ character_assets
โ โโโ orangeRobot
โ โโโ input_list
โ โ โโโ ar_cond.list
โ โ โโโ character.list
โ โ โโโ output_video.list
โ โ โโโ pos_prompt.list
โ โ โโโ scene.list
โ โ โโโ scene_mask.list
โ โโโ multi_view
โ โโโ 000_0001.png
โ โโโ 002_0001.png
โ โโโ 004_0001.png
โ โโโ 006_0001.png
โโโ scene_assets
โโโ futureUtopia
โโโ all_frame_mask.mp4
โโโ videos
โโโ 0.mp4
โโโ 1.mp4
โโโ ...
```
### Multi-GPU Inference (Recommended)
```bash
# 720P video inference
# Tested on 8 NVIDIA A100-80G GPUs
bash inference_multi_gpu.sh
```
### Single-GPU Inference
```bash
# 360P video inference
# Tested on 1 NVIDIA A100-80G GPU
bash inference_single_gpu.sh
```
### Video Merging
```bash
# Merge videos generated by auto-regressive inference
python video_merge.py --input_dir "output/orangeRobot_futureUtopia"
```
## ๐ฎ Citation
```bibtex
@misc{wang2026customx,
title={CustomX: Unified Character, Action, and Scene Customization in Video World Models},
author={Yitong Wang and Fangyun Wei and Hongyang Zhang and Bo Dai and Yan Lu},
year={2026},
eprint={2512.17796},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.17796},
}
```
## ๐ Acknowledgements
- [HunyuanCustom](https://github.com/Tencent-Hunyuan/HunyuanCustom)
- [Diffusers](https://github.com/huggingface/diffusers)
- [Transformers](https://github.com/huggingface/transformers)
- [DeepSpeed](https://github.com/deepspeedai/DeepSpeed)
- [Grand Theft Auto V](https://www.rockstargames.com/gta-v)
## ๐ License
This project is licensed under the [CC BY-SA 4.0 License](http://creativecommons.org/licenses/by-sa/4.0/). |