English
File size: 5,854 Bytes
cf20d13
 
0986e10
 
 
 
cf20d13
0986e10
 
 
 
 
 
 
 
 
 
 
 
 
 
c386b53
0986e10
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
aac0f38
0986e10
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
---
license: cc-by-sa-4.0
language:
- en
base_model:
- tencent/HunyuanCustom
---
<div align="center">

<h1 align="center">
  CustomX: Unified Character, Action, and Scene Customization in Video World Models
</h1>

<h3 align="center">
  ECCV 2026
</h3>

<p align="center">
  <a href="https://arxiv.org/abs/2512.17796"><img src="https://img.shields.io/badge/arXiv-Paper-B31B1B?style=flat&labelColor=555555&logo=arxiv&logoColor=white" alt="arXiv"></a>
  <a href="https://snowflakewang.github.io/CustomX_Page/"><img src="https://img.shields.io/badge/Project-Page-4F46E5?style=flat&labelColor=555555&logo=googlechrome&logoColor=white" alt="Project Page"></a>
  <a href="https://github.com/snowflakewang/CustomX"><img src="https://img.shields.io/badge/Code-GitHub-181717?style=flat&labelColor=555555&logo=github&logoColor=white" alt="Code"></a>
  <a href="https://huggingface.co/SnowflakeWang/CustomX"><img src="https://img.shields.io/badge/Model-Weights-FFD21E?style=flat&labelColor=555555&logo=huggingface&logoColor=FFD21E" alt="Model"></a>
</p>

</div>

---

## ๐Ÿ“– Abstract

Recent advances in world models have greatly enhanced interactive environment simulation. Existing methods mainly fall into two categories: (1) static world generation models, which construct 3D environments without active agents, and (2) controllable-entity models, which allow a single entity to perform limited actions in an otherwise uncontrollable environment. In this work, we introduce CustomX, leveraging the realism and structural grounding of static world generation while extending controllable-entity models to support user-specified characters capable of performing open-ended actions. Users can provide a 3DGS scene and a character, then use natural language to direct the character to perform diverse behaviors, ranging from basic locomotion to object-centric interactions, while freely exploring the environment. CustomX synthesizes temporally coherent video clips that preserve visual fidelity with the provided scene and character, formulated as a conditional autoregressive video generation problem. Built upon a pre-trained video generator, our training strategy significantly enhances motion dynamics while maintaining generalization across actions and characters. Our evaluation covers a broad range of aspects, including visual quality, character consistency, action controllability, and long-horizon coherence.

<div align="center">
  <img src="./assets/teaser_v3.jpg" width="100%" alt="CustomX Teaser"/>
</div>

## ๐Ÿ› ๏ธ Installation
```bash
conda create -n customx python==3.11.9
conda activate customx

pip install uv

uv pip install torch==2.4.0 torchvision==0.19.0 torchaudio==2.4.0 --index-url https://download.pytorch.org/whl/cu124
uv pip install -r requirements.txt

wget https://github.com/Dao-AILab/flash-attention/releases/download/v2.6.3/flash_attn-2.6.3+cu123torch2.4cxx11abiFALSE-cp311-cp311-linux_x86_64.whl
uv pip install flash_attn-2.6.3+cu123torch2.4cxx11abiFALSE-cp311-cp311-linux_x86_64.whl
```

## ๐Ÿ“ฆ Download Checkpoints
### HunyuanCustom Base Model
Download HunyuanCustom base model [here](https://huggingface.co/tencent/HunyuanCustom).

Only need to download the following files from HunyuanCustom:
```shell
CustomX
โ””โ”€โ”€ models
    โ””โ”€โ”€ base
        โ”œโ”€โ”€ hunyuancustom_editing_720P
        โ”‚   โ””โ”€โ”€ mp_rank_00_model_states.pt
        โ”œโ”€โ”€ vae_3d
        โ”œโ”€โ”€ openai_clip-vit-large-patch14
        โ””โ”€โ”€ llava-llama-3-8b-v1_1
```

### CustomX LoRA
```bash
TOKEN=xxx # add your Hugging Face token here

hf download --token $TOKEN \
    SnowflakeWang/CustomX \
    --local-dir models/lora
```

## ๐Ÿš€ Inference

### Input Preparation
We provide a sample input in the [input](https://github.com/snowflakewang/CustomX/tree/main/input) directory. To use your own inputs, organize them following the same structure.

```shell
input
โ”œโ”€โ”€ character_assets
โ”‚   โ””โ”€โ”€ orangeRobot
โ”‚       โ”œโ”€โ”€ input_list
โ”‚       โ”‚   โ”œโ”€โ”€ ar_cond.list
โ”‚       โ”‚   โ”œโ”€โ”€ character.list
โ”‚       โ”‚   โ”œโ”€โ”€ output_video.list
โ”‚       โ”‚   โ”œโ”€โ”€ pos_prompt.list
โ”‚       โ”‚   โ”œโ”€โ”€ scene.list
โ”‚       โ”‚   โ””โ”€โ”€ scene_mask.list
โ”‚       โ””โ”€โ”€ multi_view
โ”‚           โ”œโ”€โ”€ 000_0001.png
โ”‚           โ”œโ”€โ”€ 002_0001.png
โ”‚           โ”œโ”€โ”€ 004_0001.png
โ”‚           โ””โ”€โ”€ 006_0001.png
โ””โ”€โ”€ scene_assets
    โ””โ”€โ”€ futureUtopia
        โ”œโ”€โ”€ all_frame_mask.mp4
        โ””โ”€โ”€ videos
            โ”œโ”€โ”€ 0.mp4
            โ”œโ”€โ”€ 1.mp4
            โ””โ”€โ”€ ...
```

### Multi-GPU Inference (Recommended)
```bash
# 720P video inference
# Tested on 8 NVIDIA A100-80G GPUs
bash inference_multi_gpu.sh
```

### Single-GPU Inference
```bash
# 360P video inference
# Tested on 1 NVIDIA A100-80G GPU
bash inference_single_gpu.sh
```

### Video Merging
```bash
# Merge videos generated by auto-regressive inference
python video_merge.py --input_dir "output/orangeRobot_futureUtopia"
```

## ๐Ÿ”ฎ Citation

```bibtex
@misc{wang2026customx,
      title={CustomX: Unified Character, Action, and Scene Customization in Video World Models}, 
      author={Yitong Wang and Fangyun Wei and Hongyang Zhang and Bo Dai and Yan Lu},
      year={2026},
      eprint={2512.17796},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2512.17796}, 
}
```

## ๐Ÿ’ Acknowledgements

- [HunyuanCustom](https://github.com/Tencent-Hunyuan/HunyuanCustom)
- [Diffusers](https://github.com/huggingface/diffusers)
- [Transformers](https://github.com/huggingface/transformers)
- [DeepSpeed](https://github.com/deepspeedai/DeepSpeed)
- [Grand Theft Auto V](https://www.rockstargames.com/gta-v)

## ๐Ÿ“„ License

This project is licensed under the [CC BY-SA 4.0 License](http://creativecommons.org/licenses/by-sa/4.0/).