ViTeX-Edit-14B

馃寪 Project page  路  馃搳 Dataset  路  馃И Code  路  馃 Model weights  路  馃弳 Leaderboard

Open reference model for video scene text editing. ViTeX-Edit-14B fine-tunes the VACE branch of Wan2.1-VACE-14B on the 230-clip training split of ViTeX-Dataset and adds a glyph-video stream: the target string, rendered in a typeface matched to the source and warped along the tracked source text, is pooled into 64 tokens that every VACE block attends to. It replaces the masked scene text in a video while preserving font, color, stroke, shadow, perspective, and the surrounding scene.

Accepted to NeurIPS 2026 E&D Track.

Authors: Xinghao Chen, Xiangbo Gao, Jiongze Yu, Yuheng Wu, Zhengzhong Tu (TACO Group, Texas A&M University)

Files

file size contents
vitex_14b.safetensors 8 GB the trained VACE branch: 8 VACE blocks, glyph encoder, per-block condition cross-attention (4.02 B parameters, bf16)
base_model/ 75 GB Wan2.1-VACE-14B (DiT, uMT5-XXL, Wan VAE, tokenizer), byte-identical to Wan-AI/Wan2.1-VACE-14B

Usage

Code lives on GitHub in taco-group/ViTeX-Bench/vitex_edit: glyph-video rendering, inference (single clip or a whole split, multi-GPU, low-memory modes), the Composite post-processing wrapper, and the training recipe.

git clone https://github.com/taco-group/ViTeX-Bench.git && cd ViTeX-Bench
pip install -r vitex_edit/requirements.txt
hf download ViTeX-Bench/ViTeX-Edit-14B --local-dir models/ViTeX-Edit-14B
# already have Wan2.1-VACE-14B? download only vitex_14b.safetensors and pass --base_dir / --adapter

python vitex_edit/render_glyph.py --video source.mp4 --mask mask.mp4 \
    --source_text "HOTEL" --target_text "HILTON" --output glyph.mp4
python vitex_edit/inference.py --model_dir models/ViTeX-Edit-14B \
    --video source.mp4 --mask mask.mp4 --glyph glyph.mp4 --target_text "HILTON" --output edited.mp4

See the vitex_edit README for typeface selection, evaluation-split runs, memory options and training.

Note on the earlier inference script

This repository used to bundle an inference_example.py. It re-initialized the glyph encoder and zeroed the condition cross-attention output while loading the weights, so its outputs ignored the glyph video. That script has been removed. Please use the code on GitHub, which loads the checkpoint as trained.

License

Apache-2.0 (adapter weights). base_model/ is Wan2.1-VACE-14B under its own license (base_model/LICENSE.txt).

Citation

@inproceedings{chen2026vitexbench,
  title     = {ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing},
  author    = {Chen, Xinghao and Gao, Xiangbo and Yu, Jiongze and Wu, Yuheng and Tu, Zhengzhong},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Track on Evaluations and Datasets},
  year      = {2026},
  url       = {https://vitex-bench.github.io/}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for ViTeX-Bench/ViTeX-Edit-14B

Finetuned
(6)
this model

Space using ViTeX-Bench/ViTeX-Edit-14B 1