Instructions to use bentay85/ditto-talkinghead-trt-windows with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- TensorRT
How to use bentay85/ditto-talkinghead-trt-windows with TensorRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Ditto TalkingHead — Windows TensorRT Engines
This is a prebuilt, Windows/x86_64 TensorRT engine bundle for inference with
Ditto TalkingHead. It contains the 12
hardware-compatibility (Ampere_Plus) TRT engines we compiled from the upstream ONNX models.
Upstream / original model & code (please cite and credit these):
- 🤗 Model repo:
digital-avatar/ditto-talkinghead - 🐙 Code repo:
antgroup/ditto-talkinghead - Paper: arXiv:2411.19509 — accepted at ACM MM 2025.
The ONNX sources, PyTorch weights and the Ampere engine set are NOT shipped here.
Engines (12)
| engine | precision | role |
|---|---|---|
appearance_extractor_fp16.engine |
fp16 | appearance extractor |
blaze_face_fp16.engine |
fp16 | BlazeFace detector |
decoder_fp16.engine |
fp16 | decoder (warp→image) |
face_mesh_fp16.engine |
fp16 | face mesh (TRT port of face_mesh) |
hubert_fp32.engine |
fp32 | hubert audio encoder |
insightface_det_fp16.engine |
fp16 | insightface detector |
landmark106_fp16.engine |
fp16 | 106-point landmark |
landmark203_fp16.engine |
fp16 | 203-point landmark |
lmdm_v0.4_hubert_fp32.engine |
fp32 | LMDM v0.4 hubert decoder (diffusion) |
motion_extractor_fp32.engine |
fp32 | motion extractor |
stitch_network_fp16.engine |
fp16 | stitch network |
warp_network_fp16.engine |
fp16 | warp network (uses grid_sample_3d plugin) |
hubert, motion_extractor and lmdm_v0.4_hubert are kept fp32 deliberately — they are
precision-sensitive (audio and diffusion conditioning).
🛠️ How these engines were built (replicable)
Hardware/software used on the build machine:
| component | version |
|---|---|
| GPU | NVIDIA GeForce RTX 4060 Ti (Ampere+, sm_89) |
| Python | 3.10 |
| PyTorch | 2.5.1 (cu121) |
| TensorRT | 8.6.1 (Python wheels; engine serialization version) |
| polygraphy | 0.53.6 (used by the build script) |
| CUDA (runtime) | 12.x, driver >= 530 (Ampere+ capable) |
| grid_sample_3d plugin | embedded into warp_network_fp16.engine (grid_sample_3d_plugin DLL built from SeanWangJS/grid-sample3d-trt-plugin, with CUDA 12.8 constants + TRT 8.6.1 headers) |
The 12 engines were produced from the upstream ONNX models with the repo's converter:
uv run python old/scripts/cvt_onnx_to_trt.py --onnx_dir "checkpoints/ditto_onnx" --trt_dir "checkpoints/ditto_trt_windows"
Build settings the converter applies (see cvt_onnx_to_trt.py):
--hardware-compatibility-level=Ampere_Plus→ engines run on any Ampere or newer GPU (RTX 30/40/50 series) without rebuilding.cap[0] >= 8 ? compatible : not.--builder-optimization-level=5- Default precision: fp16 for the 9 listed above, fp32 forced for
motion_extractor,hubert,wavlmandlmdm_v0.4_hubert(precision-sensitive — we A/B testedhubert_fp16and it breaks lip-sync). warp_networkusesonnx_to_trt_for_gridsampleand embeds thegrid_sample_3d_plugininto the engine (plugins_to_serialize) — that is why the runtime never loads the plugin DLL; it is baked intowarp_network_fp16.engine. ItsGridSamplenodes are pinned to fp32 regardless of the fp16 flag.- The plugin DLL itself is the third-party
SeanWangJS/grid-sample3d-trt-plugingrid_sample_3d plugin, compiled and linked against our CUDA 12.8 toolkit and the TensorRT 8.6.1 headers, then serialized into the engine so no.dllships at runtime. - fp16 variants of the precision-sensitive models can be added with
--force_fp16 motion_extractor hubert lmdm_v0.4_hubert(kept out here becausehubert_fp16corrupted lip-sync in our tests).
These are prebuilt engines. End users only need the TensorRT 8.6.1 runtime and CUDA 12.x runtime DLLs + an Ampere-or-newer GPU — no cuDNN, ONNX runtime, mediapipe, polygraphy, CUDA toolkit (nvcc) or the plugin DLL are required for inference.
⚖️ License & Attribution
The models and code are the work of Ant Group (Ditto authors, see the paper) and are released
under Apache-2.0 (see LICENSE). When using this bundle, cite:
@article{li2024ditto,
title={Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis},
author={Li, Tianqi and Zheng, Ruobing and Yang, Minghui and Chen, Jingdong and Yang, Ming},
journal={arXiv preprint arXiv:2411.19509},
year={2024}
}
- Downloads last month
- -
Model tree for bentay85/ditto-talkinghead-trt-windows
Base model
digital-avatar/ditto-talkinghead