--- license: apache-2.0 datasets: - FastVideo/Wan-Syn_77x448x832_600k base_model: - FastVideo/FastWan2.1-T2V-1.3B-Diffusers library_name: fastvideo tags: - video-generation - text-to-video - wan - int8 - quantization - apple-silicon pipeline_tag: text-to-video --- # FastMetal-1.3B-QAD **3-step text-to-video, INT8 pre-quantized for Apple Silicon.** The entry-tier FastMetal model — a DMD2-distilled 1.3B Wan video model with a quantization-aware-trained INT8 DiT. Pre-quantized: no startup quantization, tiny download, runs on 8 GB-class Macs. ## What's inside | Path | Contents | |---|---| | `mlx_dit.safetensors` / `mlx_dit.json` | INT8 (affine, group-64) DiT | | `text_encoder/`, `vae/`, `tokenizer/`, `scheduler/` | everything needed to run standalone (fp16 UMT5 text encoder) | ## Quickstart Requires macOS with Apple silicon (MPS) and Python 3.11+: ```bash pip install torch transformers mlx safetensors av imageio imageio-ffmpeg git clone https://github.com/FastVideo/FastVideo.git cd FastVideo python examples/inference/basic/mlx_wan_prompt_to_video.py \ --model-root ./FastMetal-1.3B-QAD \ --mlx-checkpoint ./FastMetal-1.3B-QAD \ --prompt "a misty mountain river valley at sunrise" ``` ## Model details | | | |---|---| | Base model | FastWan 2.1 T2V 1.3B | | Distillation | DMD2, 3 denoising steps | | Quantization | affine INT8, group size 64, QAT-trained | | Resolution | 448×832 (480p), 77 frames | | Flow shift | 8.0 | | DiT weights | ~1.5 GB (INT8) | ## Training DMD2 distillation of the FastWan 2.1 T2V 1.3B teacher onto an INT8 student on NVIDIA GB200 clusters, with quantization-aware training (affine INT8, group 64). Training corpus: `FastVideo/Wan-Syn_77x448x832_600k`. ## FastMetal family | Model | Tier | |---|---| | [FastMetal-1.3B-QAD](https://huggingface.co/FastVideo/FastMetal-1.3B-QAD) | Entry — 16 GB+ class Macs | | [FastMetal-5B-QAD] | Mid — 720p | 16 GB+ class Macs | | [FastMetal-14B-QAD](https://huggingface.co/FastVideo/FastMetal-14B-QAD) | Quality — 24 GB+/ Ideally 36 Macs |