How to use from the
Use from the
Diffusers library
# Gated model: Login with a HF token with gated access permission
hf auth login
pip install -U diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image, export_to_video

# switch to "mps" for apple devices
pipe = DiffusionPipeline.from_pretrained("kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers", dtype=torch.bfloat16, device_map="cuda")
pipe.to("cuda")

prompt = "A man with short gray hair plays a red electric guitar."
image = load_image(
    "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png"
)

output = pipe(image=image, prompt=prompt).frames[0]
export_to_video(output, "output.mp4")

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.


KandinskyLab Report PyPI HF Demo Diffusers vLLM

ComfyUI

Kandinsky 6.0: A family of diffusion models for Video + Audio generation

We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (TI2AV) modes; a plug-in super-resolution model raises the output resolution to Full-HD (1920×1080).

Usage guide

This repository contains Kandinsky 6.0 Pro: the main Pro checkpoint, the strongest quality setting of the Pro line.

Text-to-audio-video (T2AV)

import torch
from diffusers import Kandinsky6TI2VAPipeline
from diffusers.utils import encode_video

DEFAULT_PROMPT = (
    "cinematic shot: a giant stone samurai on a stormy cliff above a neon city opens glowing golden eyes and "
    "raises a katana. Blue lightning strikes the blade, creating a massive shockwave through the clouds. The "
    "camera rapidly pulls back from a low angle. Photorealistic, epic scale, dark blue and gold lighting, rain, "
    "sparks, volumetric lightning, blockbuster quality. Audio: heavy rain, deep thunder, metallic sword hum, "
    "rising brass and choir, electrical crackle, perfectly synchronized lightning impact, sub-bass shockwave. "
    "No dialogue, text, or logos."
)

pipe = Kandinsky6TI2VAPipeline.from_pretrained(
    "kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers", torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()

output = pipe(
    prompt=DEFAULT_PROMPT,
    height=480,
    width=864,
    num_frames=121,
    num_inference_steps=50,
    guidance_scale=5.0,
)
encode_video(
    output.frames[0],
    fps=24,
    output_path="output.mp4",
    audio=output.audio[0][None],
    audio_sample_rate=pipe.audio_sample_rate,
)

Image-to-audio-video (TI2AV)

Pass a reference image with image=. It conditions the first frame, and the pipeline resizes and center-crops it to height × width. The prompt should describe what happens next in that image, so the example uses its own prompt rather than DEFAULT_PROMPT. It reuses pipe from the T2AV example.

from diffusers.utils import load_image

IMAGE_PROMPT = (
    "cinematic shot: a red dragon wakes on a mound of gold coins in a sunlit cave, slowly unfolds its wings "
    "and lowers its head toward the glowing stream below. Coins slide and clink down the pile as the dragon "
    "shifts its weight. The camera slowly pushes in toward the dragon's eyes. Photorealistic, fantasy, warm "
    "golden light, drifting dust motes, lush cave flowers. Audio: deep rumbling breath, leathery wing flaps, "
    "clinking coins, trickling water, distant echoing roar. No dialogue, text, or logos."
)

image = load_image(
    "https://huggingface.co/kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers/resolve/main/assets/i2va_input.jpg"
)

output = pipe(
    prompt=IMAGE_PROMPT,
    image=image,
    height=480,
    width=864,
    num_frames=121,
    num_inference_steps=50,
    guidance_scale=5.0,
)
encode_video(
    output.frames[0],
    fps=24,
    output_path="output_ti2av.mp4",
    audio=output.audio[0][None],
    audio_sample_rate=pipe.audio_sample_rate,
)

Options

Pass sample_audio=False to generate video only, and expand_prompts=True to let the Qwen2.5-VL text encoder rewrite a short prompt into a detailed one before encoding. expand_prompts adds latency but no extra model weights.

Super Res guide

Kandinsky 6.0 video super-resolution is a separate pipeline, Kandinsky6SRPipeline. It takes the frames produced above and upscales them by ×2, ×2.25 or ×4 with tiled diffusion in K-VAE latent space, blending overlapping tiles back with Hann windows. The example uses the flow-matching checkpoint, which runs in 4 steps per tile.

import torch
from diffusers import Kandinsky6SRPipeline
from diffusers.utils import encode_video

torch._inductor.config.max_autotune = True  # required: lets inductor pick flex-attention tiles that fit the SR block mask

sr_pipe = Kandinsky6SRPipeline.from_pretrained(
    "kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers", torch_dtype=torch.bfloat16
)
sr_pipe.transformer.set_attention_backend("flex")
sr_pipe.enable_model_cpu_offload()
sr_pipe.transformer.compile_repeated_blocks(fullgraph=True)

upscaled = sr_pipe(
    video=output.frames[0],
    resolution_scale=2.25,  # 2, 2.25 or 4
    num_inference_steps=4,  # 2 with Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers
).frames[0]
encode_video(
    upscaled,
    fps=24,
    output_path="output_sr.mp4",
    audio=output.audio[0][None],
    audio_sample_rate=pipe.audio_sample_rate,
)

The example continues from the Usage guide: pipe and output come from it. The SR pipeline works on frames only, so the audio is passed to encode_video from the generation step.

  • resolution_scale=2.25 (the default, and the route used after Kandinsky 6 generation) is a ×1.125 bilinear pre-upscale followed by the ×2 path. 2 and 4 are the other options.
  • min_overlap (default 0.2) sets the minimum tile overlap. tiles_batch_size (default 1) sets how many tiles go through the model per call, trading VRAM for speed.

Components

Folder Class Role
transformer/ Kandinsky6Transformer3DModel Multimodal video+audio diffusion transformer (Pro)
vae/ AutoencoderKLHunyuanVideo Video VAE: encodes the reference image, decodes the generated video
text_encoder/ + tokenizer/ Qwen2_5_VLForConditionalGeneration + Qwen2_5_VLProcessor Token-level text embeddings; also used for expand_prompts=True
text_encoder_2/ + tokenizer_2/ CLIPTextModel + CLIPTokenizer Pooled text embedding
audio_vae/ MMAudioVAE Audio latents ↔ mel spectrogram
vocoder/ MMAudioVocoder (BigVGAN-style) Mel spectrogram → waveform
scheduler/ FlowMatchEulerDiscreteScheduler Flow matching, shift = 5.0

model_index.json names the pipeline class Kandinsky6TI2VAPipeline.

Kandinsky 6.0 family

Checkpoint Repository Use
Pro (this repository) kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers generation
Pro, distilled, 10 steps kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers generation
Pro, pretrained kandinskylab/Kandinsky-6.0-Pro-pretrain-5s-Diffusers generation
Lite kandinskylab/Kandinsky-6.0-Lite-5s-Diffusers generation
Lite, distilled, 10 steps kandinskylab/Kandinsky-6.0-Lite-distill-5s-Diffusers generation
Lite, pretrained kandinskylab/Kandinsky-6.0-Lite-pretrain-5s-Diffusers generation
VSR, flow matching, 4 steps/tile kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers upscale
VSR, distilled, 2 steps/tile kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers upscale
Downloads last month
173
Safetensors
Model size
30B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers 1

Collection including kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers

Paper for kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers