BBInfinitetalk / backup /RESTORE.md
TechCarbasa's picture
Upload folder using huggingface_hub
bdce33b verified
|
Raw
History Blame Contribute Delete
2.3 kB

Restore InfiniteTalk + FP8 TensorRT on a fresh H100

1. Code

tar xzf code.tar.gz -C /workspace

2. Weights (NOT in this backup - re-download, they are public)

hf download Wan-AI/Wan2.1-I2V-14B-480P --local-dir /workspace/weights/Wan2.1-I2V-14B-480P
hf download MeiGen-AI/InfiniteTalk --include "single/*" --local-dir /workspace/weights/InfiniteTalk
hf download TencentGameMate/chinese-wav2vec2-base --local-dir /workspace/weights/chinese-wav2vec2-base
hf download TencentGameMate/chinese-wav2vec2-base --revision refs/pr/1 --include "model.safetensors" --local-dir /workspace/weights/chinese-wav2vec2-base
hf download Kijai/WanVideo_comfy Lightx2v/lightx2v_I2V_14B_480p_cfg_step_distill_rank64_bf16.safetensors --local-dir /workspace/weights/loras

3. Python env (pin torch, or the CUDA stack gets swapped out from under you)

uv pip install -c constraints.txt -r pip-freeze.txt

Key pins: torch==2.8.0+cu128, xformers==0.0.32.post2, transformers==4.51.3, flash_attn 2.8.3.post1 (cu12torch2.8, cxx11abiTRUE), and export XFORMERS_IGNORE_FLASH_VERSION_CHECK=1

4. TRT engines

The engines in engines/ are a CACHE, valid only on the same GPU model + TRT version (see VERSIONS.txt). Try them first:

python -c "
import tensorrt as trt
rt = trt.Runtime(trt.Logger(trt.Logger.ERROR))
e = rt.deserialize_cuda_engine(open('/workspace/engines/block0.plan','rb').read())
print('engines OK' if e else 'ENGINES INVALID -> rebuild')"

If invalid (different GPU/TRT), rebuild - 18 min total, fully scripted: # a) capture calibration activations (5 min) CAPTURE=1 <run generate_infinitetalk.py once, 81-frame clip> # see notes below # b) build 40 FP8 engines (~13 min) python /workspace/trt/build_engines.py

5. Run

/workspace/start_server.sh      # warm server, loads model once
/workspace/generate.sh <image> <audio> "<prompt>" <name>

Gotchas

  • Engines are STATIC: 81-frame windows @ 448x832 (30,576 tokens) with the 4-step I2V LoRA baked in. Change resolution / frame count / LoRA / steps -> rebuild.
  • Under WAN_TRT=1 the PyTorch blocks are DELETED (multitalk.py:689 would otherwise drag 27GB back onto the GPU).
  • All 40 engine contexts share ONE scratch buffer (5.85GB); 40 private ones would need 234GB.