--- title: Wan2.2-S2V emoji: 🗣️ colorFrom: indigo colorTo: blue sdk: static pinned: false license: apache-2.0 short_description: Reference for Wan2.2-S2V, audio-driven video generation --- # Wan2.2-S2V A reference page for **Wan2.2-S2V**, the speech-to-video checkpoint in Alibaba Tongyi Lab's Wan 2.2 family. It animates a single reference image from an audio track, driving lip sync, head motion and gesture together. - Weights: [Wan-AI/Wan2.2-S2V-14B](https://huggingface.co/Wan-AI/Wan2.2-S2V-14B) - Code: [Wan-Video/Wan2.2](https://github.com/Wan-Video/Wan2.2) - Hosted inference: [wan-2.2/speech-to-video](https://wavespeed.ai/models/wavespeed-ai/wan-2.2/speech-to-video) Wan2.2-S2V is developed and released by Alibaba Tongyi Lab. This Space is maintained by WaveSpeed AI, which provides hosted inference for the model.