Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| .gitattributes | 1.52 kB xet | 818ba6de | |
| README.md | 3.01 kB xet | 5d906b98 | |
| clip_vision_h.safetensors | 1.26 GB xet | badc2656 | |
| fantasytalking_fp16.safetensors | 1.68 GB xet | f52a4b35 | |
| umt5-xxl-enc-bf16.safetensors | 11.4 GB xet | afe10de7 | |
| wan2.1_i2v_720p_14B_fp8_e4m3fn.safetensors | 16.4 GB xet | d5c69165 |
π WanVideo Model Suite
Combined & Quantized Models for ComfyUI Workflows
Derived from Wan-AI/Wan2.1-VACE-14B
π Overview
This repository provides optimized models for WanVideoβa high-fidelity video generation framework. Models are quantized to balance performance and resource efficiency while retaining visual quality. Designed for seamless integration with ComfyUI via:
- WanVideo Wrapper (Third-party extension)
- Native WanVideo nodes in ComfyUI
π§ Key Components
1. Core Diffusion Models
| File | Size | Description |
|---|---|---|
wan2.1_i2v_720p_14B_fp8_e4m3fn.safetensors |
Quantized (FP8) | Base video generation model (14B params, 720p). |
fantasytalking_fp16.safetensors |
FP16 | Specialized model for expressive dialogue animation. |
2. Text & Vision Encoders
| File | Type | Role |
|---|---|---|
umt5-xxl-enc-bf16.safetensors |
Text Encoder (UMT5-XXL) | BF16 precision for multilingual text understanding. |
clip_vision_h.safetensors |
Vision Encoder | Processes visual inputs for conditional generation. |
π ComfyUI Setup Guide
Place files in these directories within your ComfyUI installation:
models/
βββ diffusion_models/
β βββ wan2.1_i2v_720p_14B_fp8_e4m3fn.safetensors
β βββ fantasytalking_fp16.safetensors
βββ clip_vision/
β βββ clip_vision_h.safetensors
βββ text_encoders/
βββ umt5-xxl-enc-bf16.safetensors
π Dependencies & Resources
Vision Encoder Resources
- Download
clip_vision_h.safetensorsfrom:
Comfy-Org/Wan_2.1_ComfyUI_repackaged
- Download
FantasyTalking Model
- Source code & usage: GitHub Repository
Base Model
- Full precision version: Wan-AI/Wan2.1-VACE-14B
π‘ Usage Notes
- Quantization Benefits: FP8 reduces VRAM usage by ~50% vs FP16, enabling 720p generation on consumer GPUs.
- Workflow Compatibility: Combine with
Text-to-Video,Image-to-Video, orFantasyTalkingnodes in ComfyUI. - Multi-Modal Inputs: UMT5-XXL encoder supports multilingual prompts (e.g., English, Chinese).
βοΈ License
Inherited from parent models (Check Wan-AI License). Non-commercial/research use recommended pending verification.
β¨ Pro Tip: For optimal results, pair with WanVideoβs temporal consistency modules to reduce frame flickering in long sequences.
Model Card curated by the ComfyUI community. Maintained for reproducibility and ease of deployment.
- Total size
- 30.7 GB
- Files
- 6
- Last updated
- Jul 25
- Pre-warmed CDN
- US EU US EU