Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation
Abstract
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920times1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.
Community
- π¬ Kandinsky 6.0 Video β 3B Lite / 29B Pro for synchronized text/image-to-audio-video, with 44 kHz audio, lip-sync and Full-HD super-resolution.
- π§ Uses a dual-stream CrossDiT with bidirectional audioβvideo attention, followed by SFT, RL and 10-step distillation.
- π RL post-training reduces Pro's speech WER by 47%.
- π VABench: Kandinsky 6.0 Pro achieves the best results among LTX 2.5, Kandinsky Lite, and Kandinsky Pro on speech quality, audio aesthetics, AV alignment, lip-sync, desynchronization, and visual realism.
- βοΈ Human evaluation: Kandinsky 6.0 Pro clearly outperforms Kandinsky 5.0 Pro and is preferred over LTX 2.5 for visual quality, motion realism, visual prompt following, artifact reduction, and overall task solving. It also has a statistically significant advantage over LTX 2.5 in speech quality.
- π MIT licensed, with weights, code and Diffusers integration released.
π Paper Β· GitHub Β· Hugging Face Β· Kandinsky Lab
- π¬ Kandinsky 6.0 Video Super-Resolution β a 1.4B-parameter SR diffusion transformer for Γ2, Γ2.25 and Γ4 video upscaling.
- π§ Operates in KVAE latent space: encode the input video β upscale its latents β refine them with tiled diffusion β decode the high-resolution video.
- π Available as a flow-matching model and a Ο-Flow DX distilled model requiring just 2 model evaluations per tile.
- π MIT licensed, with code and model weights released.
π Paper Β· GitHub Β· Hugging Face Β· Demo Β· Kandinsky Lab
Models citing this paper 8
kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 2
Collections including this paper 0
No Collection including this paper