ViTeX-Edit-14B
馃寪 Project page 路 馃搳 Dataset 路 馃И Code 路 馃 Model weights 路 馃弳 Leaderboard
Open reference model for video scene text editing. ViTeX-Edit-14B fine-tunes the VACE branch of Wan2.1-VACE-14B on the 230-clip training split of ViTeX-Dataset and adds a glyph-video stream: the target string, rendered in a typeface matched to the source and warped along the tracked source text, is pooled into 64 tokens that every VACE block attends to. It replaces the masked scene text in a video while preserving font, color, stroke, shadow, perspective, and the surrounding scene.
Accepted to NeurIPS 2026 E&D Track.
Authors: Xinghao Chen, Xiangbo Gao, Jiongze Yu, Yuheng Wu, Zhengzhong Tu (TACO Group, Texas A&M University)
Files
| file | size | contents |
|---|---|---|
vitex_14b.safetensors |
8 GB | the trained VACE branch: 8 VACE blocks, glyph encoder, per-block condition cross-attention (4.02 B parameters, bf16) |
base_model/ |
75 GB | Wan2.1-VACE-14B (DiT, uMT5-XXL, Wan VAE, tokenizer), byte-identical to Wan-AI/Wan2.1-VACE-14B |
Usage
Code lives on GitHub in taco-group/ViTeX-Bench/vitex_edit: glyph-video rendering, inference (single clip or a whole split, multi-GPU, low-memory modes), the Composite post-processing wrapper, and the training recipe.
git clone https://github.com/taco-group/ViTeX-Bench.git && cd ViTeX-Bench
pip install -r vitex_edit/requirements.txt
hf download ViTeX-Bench/ViTeX-Edit-14B --local-dir models/ViTeX-Edit-14B
# already have Wan2.1-VACE-14B? download only vitex_14b.safetensors and pass --base_dir / --adapter
python vitex_edit/render_glyph.py --video source.mp4 --mask mask.mp4 \
--source_text "HOTEL" --target_text "HILTON" --output glyph.mp4
python vitex_edit/inference.py --model_dir models/ViTeX-Edit-14B \
--video source.mp4 --mask mask.mp4 --glyph glyph.mp4 --target_text "HILTON" --output edited.mp4
See the vitex_edit README for typeface selection, evaluation-split runs, memory options and training.
Note on the earlier inference script
This repository used to bundle an inference_example.py. It re-initialized the glyph encoder and zeroed the condition cross-attention output while loading the weights, so its outputs ignored the glyph video. That script has been removed. Please use the code on GitHub, which loads the checkpoint as trained.
License
Apache-2.0 (adapter weights). base_model/ is Wan2.1-VACE-14B under its own license (base_model/LICENSE.txt).
Citation
@inproceedings{chen2026vitexbench,
title = {ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing},
author = {Chen, Xinghao and Gao, Xiangbo and Yu, Jiongze and Wu, Yuheng and Tu, Zhengzhong},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Track on Evaluations and Datasets},
year = {2026},
url = {https://vitex-bench.github.io/}
}
Model tree for ViTeX-Bench/ViTeX-Edit-14B
Base model
Wan-AI/Wan2.1-VACE-14B