Papers
arxiv:2610.05608

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Published on Oct 4
Β· Submitted by
Viacheslav Vasilev
on Oct 6
#1 Paper of the day
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920times1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.

Community

Paper submitter
  • 🎬 Kandinsky 6.0 Video β€” 3B Lite / 29B Pro for synchronized text/image-to-audio-video, with 44 kHz audio, lip-sync and Full-HD super-resolution.
  • 🧠 Uses a dual-stream CrossDiT with bidirectional audio↔video attention, followed by SFT, RL and 10-step distillation.
  • πŸš€ RL post-training reduces Pro's speech WER by 47%.
  • πŸ“Š VABench: Kandinsky 6.0 Pro achieves the best results among LTX 2.5, Kandinsky Lite, and Kandinsky Pro on speech quality, audio aesthetics, AV alignment, lip-sync, desynchronization, and visual realism.
  • βš”οΈ Human evaluation: Kandinsky 6.0 Pro clearly outperforms Kandinsky 5.0 Pro and is preferred over LTX 2.5 for visual quality, motion realism, visual prompt following, artifact reduction, and overall task solving. It also has a statistically significant advantage over LTX 2.5 in speech quality.
  • πŸ”“ MIT licensed, with weights, code and Diffusers integration released.

πŸ”— Paper Β· GitHub Β· Hugging Face Β· Kandinsky Lab

  • 🎬 Kandinsky 6.0 Video Super-Resolution β€” a 1.4B-parameter SR diffusion transformer for Γ—2, Γ—2.25 and Γ—4 video upscaling.
  • 🧠 Operates in KVAE latent space: encode the input video β†’ upscale its latents β†’ refine them with tiled diffusion β†’ decode the high-resolution video.
  • πŸš€ Available as a flow-matching model and a Ο€-Flow DX distilled model requiring just 2 model evaluations per tile.
  • πŸ”“ MIT licensed, with code and model weights released.

πŸ”— Paper Β· GitHub Β· Hugging Face Β· Demo Β· Kandinsky Lab

Sign up or log in to comment

Models citing this paper 8

Browse 8 models citing this paper

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.05608 in a dataset README.md to link it from this page.

Spaces citing this paper 2

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.