Kandinsky 6.0: ComfyUI Setup and Model Guide

ComfyUI Wiki

Run Kandinsky 6.0 in ComfyUI: 29B Pro and 3B Lite generate 5-second video with synchronized 44 kHz audio, plus a super-resolution model for Full HD.

K

Kandinsky 6.0

Video GenerationAudio-VideoText-to-VideoImage-to-VideoComfyUI

Kandinsky 6.0 Video is a family of foundation diffusion models for synchronized text-to-audio-video generation. Pro (29B) and Lite (3B) both produce 5-second clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video and image-to-audio-video modes, and a separate super-resolution model raises the output to Full HD.

DeveloperKandinsky Lab
Release Date2026-10-06 (open weights)
ArchitectureMultimodal video and audio diffusion transformer with flow matching (29B Pro, 3B Lite)
Text EncoderQwen2.5-VL with a CLIP pooled embedding
AudioMMAudio VAE and BigVGAN-style vocoder, synchronized 44 kHz output
Output5-second clips at 24 fps, up to Full HD after super resolution
Generation ModesText-to-audio-video, image-to-audio-video
LicenseMIT

Kandinsky 6.0 Video generates image and soundtrack together in one pass rather than dubbing audio on afterwards. A multimodal video and audio diffusion transformer models both streams, with a Qwen2.5-VL text encoder supplying token-level text embeddings and a CLIP model supplying a pooled embedding. The visual side decodes through a Hunyuan Video VAE, and the audio side runs an MMAudio VAE plus a BigVGAN-style vocoder for 44 kHz waveform output. Sampling uses flow matching with shift = 5.0.

Models

The family ships in two sizes, each with a base, a distilled 10-step and a pretrained checkpoint, plus a dedicated super-resolution pair:

CheckpointSizeNotes
Kandinsky 6.0 Pro29BFlagship generation checkpoint
Kandinsky 6.0 Pro distill29BDistilled to 10 steps with PiFlow, used by default in the ComfyUI templates
Kandinsky 6.0 Pro pretrain29BPretrained base for fine-tuning
Kandinsky 6.0 Lite3BCompact generation checkpoint
Kandinsky 6.0 Lite distill3BDistilled 10-step variant
Kandinsky 6.0 Lite pretrain3BPretrained base for fine-tuning
Kandinsky 6.0 VSR-Super-resolution pipeline, x2, x2.25 and x4, 4 steps per tile
Kandinsky 6.0 VSR distill-Distilled super-resolution, 2 steps per tile

Super resolution

Kandinsky6SRPipeline upscales generated frames with tiled diffusion in K-VAE latent space and blends overlapping tiles with Hann windows. resolution_scale accepts 2, 2.25 (the default route after generation) or 4, and min_overlap and tiles_batch_size trade tile blending and VRAM for speed.

ComfyUI

The official ComfyUI extension is published by kandinskylab on the ComfyUI Registry as kandinsky6 and kandinsky6-sr. Install both through ComfyUI Manager, restart, and open Workflow → Browse Templates → kandinsky6 for the bundled Text to Video+Audio and Image to Video+Audio templates. Model files, including the companion Diffusers JSON configs, are fetched by the extension's Download models button into the correct ComfyUI folders.

Resources

Guides and workflows related to this model series.

No articles found.

Comments

Sign in with GitHub to join the discussion.

Loading comments…