Text-to-Video
Diffusers
Safetensors
MiniMaxH3ModularPipeline
video-generation
audio
fastvideo
bf16
pruned
Instructions to use FastVideo/FastVideo-FastH3-Trim-8-Step with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use FastVideo/FastVideo-FastH3-Trim-8-Step with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("FastVideo/FastVideo-FastH3-Trim-8-Step", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
|
Download README.md from FastVideo/FastVideo-FastH3-Trim-8-Step: direct link, hf CLI and curl.
- Browser
- Download file 1.54 kB
-
https://huggingface.co/FastVideo/FastVideo-FastH3-Trim-8-Step/resolve/main/README.md
- Command line
-
hf download hf://FastVideo/FastVideo-FastH3-Trim-8-Step/README.md
-
curl -L -o README.md https://huggingface.co/FastVideo/FastVideo-FastH3-Trim-8-Step/resolve/main/README.md
1.54 kB
| license: other | |
| base_model: FastVideo/FastVideo-FastH3-8-Step-V2 | |
| tags: [video-generation, text-to-video, audio, fastvideo, bf16, pruned] | |
| # FastH3 Trim, 8-step, BF16 source weights | |
| FastH3 Trim is an experimental, smaller version of [FastH3 V2](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2): | |
| eight-step text-to-video with synchronized audio from 42 of the 50 MiniMax H3 transformer blocks. We removed the eight | |
| blocks whose removal changed video and audio predictions the least, replaced each block's AdaLN timestep projection with | |
| a shared rank-16 basis, and trained the result with eight-step DMD2 (checkpoint 300). Removing blocks makes the model | |
| faster and smaller but costs some quality; use FastH3 V2 when quality matters most. | |
| - **Sampling:** 8 DMD steps (999, 874, 749, 624, 500, 375, 250, 125), sparse attention keeping 20% of tiles. | |
| `fastvideo_inference.json` holds the schedule, which FastVideo reads automatically. | |
| - **Text encoder:** NVFP4 Qwen3-VL trimmed to the 50 layers H3 reads. | |
| - **VAE:** LynnReal lightweight video VAE with Kijai's INT8 weights; H3 audio VAE. | |
| Blog post: [FastH3 on Consumer Hardware](https://haoailab.com/blogs/fasth3-rtx/) · Code: [FastVideo](https://github.com/hao-ai-lab/FastVideo) | |
| - **Transformer:** 34.9 GiB in BF16 (FP16 AdaLN factors). Source for the quantized releases and for local MLX conversion on Apple Silicon. | |
| Full-quality counterpart in the same format: [FastVideo/FastVideo-FastH3-8-Step-V2](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2). | |